bigbio/ebm_pico
收藏资源简介:
该语料库包含4,993篇摘要,标注了参与者、干预措施和结果。训练标签来自AMT工作者,测试标签来自医疗专业人员。
This corpus contains 4,993 abstracts annotated with participants, interventions, and outcomes. Training labels are sourced from AMT workers, while test labels are provided by medical professionals.
数据集概述
基本信息
- 语言: 英语
- 多语言性: 单语
- 许可证: 未知
- 任务: 命名实体识别(NER)
数据集详情
- 主页: https://github.com/bepnye/EBM-NLP
- 是否公开: 是
- 是否可用于PubMed: 是
- 数据内容: 包含4,993篇经过标注的摘要,标注内容包括参与者(P)、干预措施(I)和结果(O)。训练标签由AMT工作者提供并经过聚合以减少噪声,测试标签由医学专业人员收集。
引用信息
@inproceedings{nye-etal-2018-corpus, title = "A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature", author = "Nye, Benjamin and Li, Junyi Jessy and Patel, Roma and Yang, Yinfei and Marshall, Iain and Nenkova, Ani and Wallace, Byron", booktitle = "Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)", month = jul, year = "2018", address = "Melbourne, Australia", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/P18-1019", doi = "10.18653/v1/P18-1019", pages = "197--207", }




