遇见数据集

Annotated dataset mentions corpus in IR/ML/NLP domain

收藏
Zenodo2023-06-13 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

This corpus is a re-annotated version of the dataset available at https://github.com/xjaeh/ner_dataset_recognition and described in the following publication: Heddes, J.; Meerdink, P.; Pieters, M.; Marx, M. The Automatic Detection of Dataset Names in Scientific Articles. <em>Data</em> 2021, <em>6</em>, 84. https://doi.org/10.3390/data6080084 The corpus contains 6000 sentences in the IR/ML/NLP domains, with dataset annotations. The original corpus in CSV covered only explicitly named and reused datasets. In addition, "conjunctions" of datasets were annotated in a single span. We review entirely the annotation to include new datasets too (as developed in the described research work of the source articles) and to annotate separately every individual datasets. In addition, we re-packaged the corpus into a more standard JSON with annotation offsets. Python scripts for conversion are available at https://github.com/kermitt2/dataset_recognition_resources We thank the original authors of the corpus for their very valuable resource !

本语料库为https://github.com/xjaeh/ner_dataset_recognition 公开的数据集的重注释版本,相关研究发表于下述论文:Heddes, J.; Meerdink, P.; Pieters, M.; Marx, M. *The Automatic Detection of Dataset Names in Scientific Articles*(《科学文章中数据集名称的自动检测》),*Data* 2021, *6*, 84. https://doi.org/10.3390/data6080084。该语料库包含信息检索(Information Retrieval, IR)、机器学习(Machine Learning, ML)、自然语言处理(Natural Language Processing, NLP)领域的6000条带数据集标注的句子。原始CSV格式语料库仅涵盖显式命名且被复用的数据集,此外还对单个标注区间内的数据集联合项进行了注释。我们对原有注释进行了全面修订,不仅新增了源论文所述研究工作中开发的新数据集,还将每个独立数据集单独标注。此外,我们将该语料库重新封装为带有标注偏移量的标准JSON格式。用于格式转换的Python脚本可从https://github.com/kermitt2/dataset_recognition_resources 获取。我们谨向该语料库的原作者致谢,感谢他们提供的宝贵研究资源!

提供机构:
Zenodo
创建时间:
2023-06-13
二维码
社区交流群
二维码
科研交流群
商业服务