bigbio/medmentions
收藏资源简介:
--- language: - en bigbio_language: - English license: cc0-1.0 multilinguality: monolingual bigbio_license_shortname: CC0_1p0 pretty_name: MedMentions homepage: https://github.com/chanzuckerberg/MedMentions bigbio_pubmed: True bigbio_public: True bigbio_tasks: - NAMED_ENTITY_DISAMBIGUATION - NAMED_ENTITY_RECOGNITION --- # Dataset Card for MedMentions ## Dataset Description - **Homepage:** https://github.com/chanzuckerberg/MedMentions - **Pubmed:** True - **Public:** True - **Tasks:** NED,NER MedMentions is a new manually annotated resource for the recognition of biomedical concepts. What distinguishes MedMentions from other annotated biomedical corpora is its size (over 4,000 abstracts and over 350,000 linked mentions), as well as the size of the concept ontology (over 3 million concepts from UMLS 2017) and its broad coverage of biomedical disciplines. Corpus: The MedMentions corpus consists of 4,392 papers (Titles and Abstracts) randomly selected from among papers released on PubMed in 2016, that were in the biomedical field, published in the English language, and had both a Title and an Abstract. Annotators: We recruited a team of professional annotators with rich experience in biomedical content curation to exhaustively annotate all UMLS® (2017AA full version) entity mentions in these papers. Annotation quality: We did not collect stringent IAA (Inter-annotator agreement) data. To gain insight on the annotation quality of MedMentions, we randomly selected eight papers from the annotated corpus, containing a total of 469 concepts. Two biologists ('Reviewer') who did not participate in the annotation task then each reviewed four papers. The agreement between Reviewers and Annotators, an estimate of the Precision of the annotations, was 97.3%. ## Citation Information ``` @misc{mohan2019medmentions, title={MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts}, author={Sunil Mohan and Donghui Li}, year={2019}, eprint={1902.09476}, archivePrefix={arXiv}, primaryClass={cs.CL} } ```
语言:英语 大生物信息学数据集语言:英语 许可证:CC0 1.0 多语言属性:单语言 大生物信息学数据集短许可证名:CC0_1p0 正式名称:MedMentions 主页:https://github.com/chanzuckerberg/MedMentions 大生物信息学数据集关联PubMed:是 大生物信息学数据集公开状态:是 大生物信息学数据集任务:命名实体消歧(NAMED_ENTITY_DISAMBIGUATION)、命名实体识别(NAMED_ENTITY_RECOGNITION) # MedMentions 数据集卡片 ## 数据集描述 - **主页:** https://github.com/chanzuckerberg/MedMentions - **关联PubMed:** 是 - **公开状态:** 是 - **任务:** 命名实体消歧(NED)、命名实体识别(NER) MedMentions 是一款全新的人工标注生物医学概念识别资源。与其他已标注生物医学语料库相比,其显著优势在于三大特征:一是体量规模庞大,包含超4000篇学术摘要与超35万个关联实体提及;二是配套的概念本体库覆盖广泛,包含源自UMLS®(统一医学语言系统,Unified Medical Language System)2017的超300万个概念;三是涵盖多类生物医学学科领域。 ### 语料库说明 MedMentions 语料库共包含4392篇论文(含标题与摘要),均为2016年发布于PubMed的生物医学领域英文文献,且同时具备完整标题与摘要,经随机抽样筛选得到。 ### 标注人员说明 我们招募了一支具备丰富生物医学内容编目经验的专业标注团队,对所有论文中的UMLS®(2017AA完整版)实体提及进行了全覆盖标注。 ### 标注质量评估 本数据集未采集严格的标注者间一致性(Inter-annotator agreement, IAA)数据。为评估MedMentions的标注质量,我们从已标注语料库中随机抽取8篇论文,共计包含469个概念。随后由两名未参与本次标注任务的生物学家担任评审员,各自评审其中4篇论文。评审员与原标注人员的一致性(即标注精度的评估指标)达到97.3%。 ## 引用信息 @misc{mohan2019medmentions, title={MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts}, author={Sunil Mohan and Donghui Li}, year={2019}, eprint={1902.09476}, archivePrefix={arXiv}, primaryClass={cs.CL} }
数据集概述:MedMentions
基本信息
- 名称: MedMentions
- 语言: 英语
- 许可证: CC0-1.0
- 多语言性: 单语
- 官方主页: https://github.com/chanzuckerberg/MedMentions
- 是否公开: 是
- 是否包含PubMed数据: 是
数据集描述
- 任务:
- 实体消歧(NAMED_ENTITY_DISAMBIGUATION)
- 实体识别(NAMED_ENTITY_RECOGNITION)
- 数据集特点:
- 包含超过4,000篇摘要和超过350,000个链接提及
- 使用超过300万个UMLS 2017概念的广泛生物医学学科覆盖
- 数据来源:
- 随机选取2016年发布在PubMed上的4,392篇生物医学领域的英文论文(标题和摘要)
- 标注者:
- 专业的生物医学内容管理团队
- 标注质量:
- 通过两位未参与标注的生物学家对8篇论文(共469个概念)的审查,标注精度达到97.3%
引用信息
@misc{mohan2019medmentions, title={MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts}, author={Sunil Mohan and Donghui Li}, year={2019}, eprint={1902.09476}, archivePrefix={arXiv}, primaryClass={cs.CL} }




