TASTEset
收藏资源简介:
TASTEset是由华沙理工大学数学与信息科学学院的研究团队开发的一个包含700条食谱的数据集,旨在为食品计算领域提供一个全面的实体识别基准。该数据集涵盖了超过13,000个实体,包括食品产品、数量及其单位、烹饪过程名称、原料的物理质量、用途和味道等。数据集的创建过程涉及从多个网站手动收集和标注食谱,使用BRAT标注工具进行实体识别。TASTEset的应用领域广泛,包括食谱相似性分析、基于深度食品知识的食谱推荐、多语言食谱翻译、新食谱生成及营养成分估算等,旨在解决食品计算中的复杂信息提取问题。
TASTEset is a dataset of 700 recipes developed by a research team from the Faculty of Mathematics and Information Science at Warsaw University of Technology, designed to serve as a comprehensive entity recognition benchmark for the field of food computing. This dataset encompasses over 13,000 entities including food products, quantities and their corresponding units, names of cooking processes, physical properties of ingredients, their intended uses and flavors, and so on. The development of TASTEset involved manual collection and annotation of recipes from multiple websites, with entity recognition performed using the BRAT annotation tool. TASTEset has a wide range of application scenarios, such as recipe similarity analysis, deep food knowledge-based recipe recommendation, multilingual recipe translation, novel recipe generation and nutritional component estimation, among others. It is intended to resolve complex information extraction issues in the field of food computing.




