遇见数据集

Experiment-I configuration details.

收藏
Figshare2025-01-28 更新2026-04-28 收录
官方服务:

资源简介:

Sharing cooking recipes is a great way to exchange culinary ideas and provide instructions for food preparation. However, categorizing raw recipes found online into appropriate food genres can be challenging due to a lack of adequate labeled data. In this study, we present a dataset named the “Assorted, Archetypal, and Annotated Two Million Extended (3A2M+) Cooking Recipe Dataset” that contains two million culinary recipes labeled in respective categories with extended named entities extracted from recipe descriptions. This collection of data includes various features such as title, NER, directions, and extended NER, as well as nine different labels representing genres including bakery, drinks, non-veg, vegetables, fast food, cereals, meals, sides, and fusions. The proposed pipeline named 3A2M+ extends the size of the Named Entity Recognition (NER) list to address missing named entities like heat, time or process from the recipe directions using two NER extraction tools. 3A2M+ dataset provides a comprehensive solution to the various challenging recipe-related tasks, including classification, named entity recognition, and recipe generation. Furthermore, we have demonstrated traditional machine learning, deep learning and pre-trained language models to classify the recipes into their corresponding genre and achieved an overall accuracy of 98.6%. Our investigation indicates that the title feature played a more significant role in classifying the genre.

分享烹饪食谱是交流烹饪创意、提供食物制作指导的绝佳途径。然而,由于缺乏充足的标注数据,将网络上获取的原始食谱归类至合适的食品品类颇具挑战。本研究推出了一款名为"多样化、典型化且带标注的200万扩展版烹饪食谱数据集(Assorted, Archetypal, and Annotated Two Million Extended (3A2M+) Cooking Recipe Dataset)"的数据集,该数据集包含200万条已按类别标注的烹饪食谱,并从食谱描述中提取了扩展命名实体。该数据集涵盖标题、命名实体识别(Named Entity Recognition,NER)结果、制作步骤以及扩展命名实体等多种特征,同时包含9种品类标签,涵盖烘焙、饮品、非素食、蔬菜类、快餐、谷物、主餐、配菜以及融合料理。本研究提出的3A2M+处理流程借助两款命名实体识别工具,对命名实体识别列表进行扩展,以弥补食谱制作步骤中缺失的热量、时长或流程等命名实体。3A2M+数据集可为各类与食谱相关的挑战性任务提供全面解决方案,涵盖分类、命名实体识别以及食谱生成等任务。此外,我们通过传统机器学习、深度学习以及预训练语言模型开展食谱品类分类实验,最终整体分类准确率达到98.6%。研究表明,标题特征在食谱品类分类任务中起到了更为关键的作用。

创建时间:
2025-01-28
二维码
社区交流群
二维码
科研交流群
商业服务