bigbio/an_em
收藏资源简介:
AnEM语料库是一个领域和物种无关的资源,手动注释了使用细粒度分类系统的解剖实体提及。该语料库包含500个文档(超过90,000字),这些文档是从引用摘要和全文论文中随机选择的,旨在代表整个可用的生物医学科学文献。语料库的注释涵盖了健康和病理解剖实体的提及,包含超过3,000个注释提及。
The AnEM corpus is a domain- and species-agnostic resource with manually annotated anatomical entity mentions using a fine-grained classification system. This corpus consists of 500 documents (over 90,000 words) randomly selected from cited abstracts and full-text articles, designed to represent the entire available biomedical scientific literature. The annotations in the corpus cover mentions of both healthy and pathological anatomical entities, with more than 3,000 annotated mentions in total.
数据集概述
基本信息
- 数据集名称: AnEM
- 语言: 英语
- 许可证: CC-BY-SA-3.0
- 多语言性: 单语种
- PubMed可用性: 是
- 公开可用性: 是
任务类型
- 命名实体识别 (NER)
- 共指消解 (COREF)
- 关系抽取 (RE)
数据集详情
- 描述: AnEM是一个领域和物种独立的手工标注资源,用于解剖实体提及,使用细粒度分类系统。该数据集包含500个文档(超过90,000字),随机选自引文摘要和全文论文,旨在代表整个可用的生物医学科学文献。
- 标注内容: 包含健康和病理解剖实体的提及,共有超过3,000个标注提及。
引用信息
@inproceedings{ohta-etal-2012-open, author = {Ohta, Tomoko and Pyysalo, Sampo and Tsujii, Jun{}ichi and Ananiadou, Sophia}, title = {Open-domain Anatomical Entity Mention Detection}, journal = {}, volume = {W12-43}, year = {2012}, url = {https://aclanthology.org/W12-4304}, doi = {}, biburl = {}, bibsource = {}, publisher = {Association for Computational Linguistics} }




