COMETA
收藏资源简介:
COMETA数据集是由剑桥大学语言技术实验室创建的,包含20,015条来自Reddit的英文生物医学实体提及,这些提及均由专家标注并与SNOMED CT知识图谱链接。数据集涵盖了从症状、疾病到化学物质、基因等多种概念,旨在解决社交媒体中健康领域实体链接的复杂性问题。创建过程中,研究人员从Reddit中筛选并爬取了高质量的健康相关讨论,通过Flair NER系统识别实体,并由专业注释者进行标注。COMETA数据集的应用领域主要集中在提升社交媒体中健康相关文本的实体链接技术,特别是在处理非正式语言和复杂医学术语时的挑战。
The COMETA dataset was developed by the Language Technology Laboratory at the University of Cambridge. It contains 20,015 English biomedical entity mentions collected from Reddit, all of which were expert-annotated and linked to the SNOMED CT knowledge graph. The dataset covers a diverse set of concepts spanning symptoms, diseases, chemical substances, genes, and more, and is designed to address the complexities of entity linking in the healthcare domain on social media. During its construction, researchers filtered and crawled high-quality health-related discussions from Reddit, identified entities using the Flair NER system, and had the entities annotated by professional annotators. The primary applications of the COMETA dataset focus on advancing entity linking technologies for health-related texts on social media, particularly overcoming the challenges posed by informal language and complex medical terminology.
- 1COMETA: A Corpus for Medical Entity Linking in the Social Media剑桥大学语言技术实验室 · 2020年



