Revised JNLPBA Corpus
收藏资源简介:
Revised JNLPBA Corpus是由信息科学研究所,中央研究院在台北创建的生物医学命名实体识别(BNER)和生物医学关系抽取(BRE)任务的数据集。该数据集保留了原始实体类型,包括蛋白质、DNA、RNA、细胞系和细胞类型,并通过领域专家根据新的标注指南重新手工整理了所有摘要。数据集创建过程中,指出了JNLPBA中的一些不完美问题并进行了修正。Revised JNLPBA Corpus主要应用于生物医学关系抽取任务,旨在提高NER系统在生物医学文本中的性能和准确性。
The Revised JNLPBA Corpus is a dataset for biomedical named entity recognition (BNER) and biomedical relation extraction (BRE) tasks, developed by the Institute of Information Science, Academia Sinica in Taipei. This dataset retains the original entity types including proteins, DNA, RNA, cell lines and cell types, and all abstracts have been manually reannotated by domain experts in accordance with new annotation guidelines. During the creation of this dataset, some imperfections in the original JNLPBA were identified and corrected. The Revised JNLPBA Corpus is primarily applied to biomedical relation extraction tasks, aiming to improve the performance and accuracy of named entity recognition systems in biomedical texts.




