EBM-NLP
收藏资源简介:
EBM-NLP数据集是由东北大学等机构创建,包含5000篇医学文章的丰富注释摘要,描述了临床随机对照试验。数据集详细标注了描述患者群体、干预措施和测量结果(PICO元素)的文本范围,并进一步在更细粒度上进行标注,如标记和映射到结构化医学词汇中的个别干预措施。数据集通过混合众包标注策略,使用不同专业水平和成本的异质标注者,从普通人群到医学博士,进行标注。该数据集旨在支持医学文献的搜索和基于证据的医学实践,解决医学干预选择的信息组织和搜索难题。
The EBM-NLP Dataset was developed by Northeastern University and other institutions. It contains 5,000 fully annotated abstracts of medical articles that describe clinical randomized controlled trials (RCTs). The dataset meticulously annotates the textual spans corresponding to patient populations, interventions, and outcome measures (collectively referred to as PICO elements), and further provides fine-grained annotations including labeling individual interventions and mapping them to structured medical vocabularies. The dataset was annotated using a hybrid crowdsourcing annotation strategy, employing heterogeneous annotators across varying professional levels and cost tiers, ranging from general members of the public to medical doctors. This dataset aims to support medical literature retrieval and evidence-based medical practice, addressing the challenges of information organization and search for medical intervention selection.
- 1A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature东北大学 · 2018年



