TDMSci
收藏资源简介:
TDMSci是由IBM研究院欧洲分部在爱尔兰创建的一个专门用于科学文献实体标注的数据集,包含从NLP论文中提取的2000个句子,由领域专家标注了任务(T)、数据集(D)和度量(M)实体。该数据集的创建旨在通过自动构建NLP领域的TDM分类法,帮助研究人员快速理解相关文献或进行可比性实验。数据集的应用领域主要集中在科学出版物摘要和知识发现,旨在解决研究人员在特定领域跟踪所有研究发表的困难,减少研究重复和基准过时的问题。
TDMSci is a specialized dataset for scientific literature entity annotation, created by IBM Research Europe in Ireland. It contains 2000 sentences extracted from NLP papers, with entities of Task (T), Dataset (D) and Metric (M) annotated by domain experts. This dataset was developed to automatically construct a TDM taxonomy for the NLP domain, helping researchers quickly understand relevant literature or conduct comparative experiments. The main application areas of this dataset focus on scientific publication abstracts and knowledge discovery, aiming to address the difficulties that researchers encounter when tracking all published research in a specific field, and reduce the problems of research duplication and outdated benchmarks.




