bigbio/nlmchem
收藏资源简介:
NLM-Chem语料库由150篇来自PubMed Central开放获取数据集的全文学术文章组成,涵盖67种不同的化学期刊,旨在覆盖生物医学文献中化学名称使用的广泛分布。文章的选择标准是那些对人类注释最有价值的文章,即富含生物实体且当前最先进的命名实体识别系统在生物实体识别上存在分歧的文章。数据集支持的任务包括命名实体识别(NER)、命名实体消歧(NED)和文本分类(TXTCLASS)。
The NLM-Chem Corpus consists of 150 full-length academic articles sourced from the PubMed Central Open Access Dataset, covering 67 distinct chemistry journals. It is designed to cover the broad distribution of chemical name usage in biomedical literature. The article selection criteria prioritize articles that are most valuable for human annotation, specifically those rich in biological entities and where state-of-the-art named entity recognition (NER) systems exhibit disagreements on biological entity recognition. Tasks supported by this dataset include named entity recognition (NER), named entity disambiguation (NED), and text classification (TXTCLASS).
数据集概述
基本信息
- 语言: 英语
- 许可证: CC0-1.0
- 多语言性: 单语种
- PubMed可用性: 是
- 公开性: 是
数据集内容
- 包含文献数量: 150篇
- 来源期刊: 67种化学期刊
- 文献来源: PubMed Central Open Access
- 数据集目的: 覆盖生物医学文献中化学名称的广泛使用,特别选择富含生物实体且现有最先进的命名实体识别系统在生物实体识别上存在分歧的文章。
任务类型
- 命名实体识别 (NER)
- 命名实体消歧 (NED)
- 文本分类 (TXTCLASS)
引用信息
@Article{islamaj2021nlm, title={NLM-Chem, a new resource for chemical entity recognition in PubMed full text literature}, author={Islamaj, Rezarta and Leaman, Robert and Kim, Sun and Kwon, Dongseop and Wei, Chih-Hsuan and Comeau, Donald C and Peng, Yifan and Cissel, David and Coss, Cathleen and Fisher, Carol and others}, journal={Scientific Data}, volume={8}, number={1}, pages={1--12}, year={2021}, publisher={Nature Publishing Group} }




