bigbio/nlm_wsd
收藏资源简介:
为了支持使用自然语言处理技术自动解决词义歧义的研究,我们构建了这个医学文本测试集,其中的歧义由人工解析。评估者被要求检查歧义词的实例,并通过选择最能代表该意义的Metathesaurus概念来确定其意图意义。该测试集包含50个1998年MEDLINE中高度频繁的UMLS概念,每个概念有100个随机选择的歧义实例,总计5000个实例。共有11名评估者参与,其中8名完成了所有5000个实例,1名完成了56%,1名完成了44%,最后一名完成了12%的实例。评估结果仅在评估者完成给定歧义的所有100个实例时使用。
To support research on automatically resolving word sense disambiguation using natural language processing technologies, we constructed this medical text test set where ambiguities are manually resolved. Annotators were asked to examine instances of ambiguous words and determine their intended meanings by selecting the Metathesaurus concept that best represents the corresponding sense. This test set contains 50 highly frequent UMLS concepts from the 1998 MEDLINE database, with 100 randomly selected ambiguous instances per concept, totaling 5,000 instances. A total of 11 annotators participated: 8 completed all 5,000 instances, 1 completed 56% of the instances, 1 completed 44%, and the last annotator completed 12% of the instances. Only the annotation results from annotators who finished all 100 instances of a given ambiguous word sense are utilized for evaluation.
数据集概述
基本信息
- 名称: NLM WSD
- 语言: 英语
- 许可证: UMLS_LICENSE
- 多语言性: 单语种
- 是否公开: 否
- 是否可在PubMed上找到: 是
数据集描述
- 任务: 命名实体消歧(NAMED_ENTITY_DISAMBIGUATION)
- 构建目的: 支持研究自动解决医学文本中的词义歧义,使用自然语言处理技术。
- 数据来源: 1998年MEDLINE文献,包含50个高频歧义的UMLS概念,每个概念有100个随机选定的歧义实例,总计5,000个实例。
- 评估者: 共11位评估者,其中8位完成全部5,000个实例,其余评估者完成比例分别为56%、44%和12%。
- 评估标准: 评估者需完成所有100个实例的评估,评估结果才被采用。
引用信息
@article{weeber2001developing, title = "Developing a test collection for biomedical word sense disambiguation", author = "Weeber, M and Mork, J G and Aronson, A R", journal = "Proc AMIA Symp", pages = "746--750", year = 2001, language = "en" }




