遇见数据集

Ontology Enrichment from Texts (OET): A Biomedical Dataset for Concept Discovery and Placement

收藏
Zenodo2023-12-26 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

A biomedical dataset supporting ontology enrichment from texts, by concept discovery and placement, adapting the MedMentions dataset (PubMed abstracts) with SNOMED CT of versions in 2014 and 2017 under the Diseases (disorder) sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic (CPP) product. The dataset is documented in the work, <em>Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and Placement</em>, on arXiv: https://arxiv.org/abs/2306.14704 (CIKM 2023). The companion code is available at https://github.com/KRR-Oxford/OET. Out-of-KB mention discovery (including the settings of mention-level data) is further partly documented in the work, <em>Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking</em>, on arXiv: https://arxiv.org/abs/2302.07189 (CIKM 2023). ver3: we revised and updated mention-level data (syn_full, synonym augmentation setting) and the folder structure, and also updated the edge catalogues with complex edges. ver2: we revised the mention-level data by only keeping out-of-KB mentions (or "NIL" mentions) associated with one-hop edges (including leaf nodes, as &lt;leaf node, NULL&gt;) and two-hop edges in the ontology (SNOMED CT 20140901). Acknowledgement of data sources and tools below: * SNOMED CT https://www.nlm.nih.gov/healthit/snomedct/archive.html (and use snomed-owl-toolkit to form .owl files)<br> * UMLS https://www.nlm.nih.gov/research/umls/licensedcontent/umlsarchives04.html (and mainly use MRCONSO for mapping UMLS to SNOMED CT)<br> * MedMentions https://github.com/chanzuckerberg/MedMentions (source of entity linking) * Protege http://protegeproject.github.io/protege/<br> * snomed-owl-toolkit https://github.com/IHTSDO/snomed-owl-toolkit<br> * DeepOnto https://github.com/KRR-Oxford/DeepOnto (based on OWLAPI https://owlapi.sourceforge.net/) for ontology processing and complex concept verbalisation

本数据集为一款面向文本本体富集 (ontology enrichment) 的生物医学数据集,通过概念发现与定位实现相关功能,以MedMentions数据集(PubMed摘要 (PubMed abstracts))为基础构建,整合了2014与2017版医学系统命名法临床术语 (SNOMED CT),涵盖疾病(紊乱)子类别以及临床发现、操作、药物/生物制品(CPP产品)三大宽泛分类。 该数据集的相关研究成果刊载于论文<em>Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and Placement</em>,收录于arXiv平台:https://arxiv.org/abs/2306.14704(CIKM 2023会议)。配套代码可从https://github.com/KRR-Oxford/OET获取。 知识库外提及发现 (Out-of-KB mention discovery,包含提及级数据的相关设置) 的相关内容另有部分刊载于论文<em>Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking</em>,收录于arXiv平台:https://arxiv.org/abs/2302.07189(CIKM 2023会议)。 版本3(ver3):我们修订并更新了提及级数据(syn_full、同义词增强设置)以及文件夹结构,同时针对复杂边更新了边目录。 版本2(ver2):我们对提及级数据进行了修订,仅保留与本体(SNOMED CT 20140901)中的单跳边(包含叶节点,形式为<leaf node, NULL>)和双跳边相关联的知识库外提及(或称"NIL提及")。 以下为数据集来源与工具的致谢声明: * 医学系统命名法临床术语 (SNOMED CT):https://www.nlm.nih.gov/healthit/snomedct/archive.html(使用snomed-owl-toolkit生成.owl格式文件) * 统一医学语言系统 (UMLS):https://www.nlm.nih.gov/research/umls/licensedcontent/umlsarchives04.html(主要使用MRCONSO表实现UMLS与SNOMED CT的映射) * MedMentions数据集:https://github.com/chanzuckerberg/MedMentions(实体链接任务的数据源) * Protege工具:http://protegeproject.github.io/protege/ * snomed-owl-toolkit工具:https://github.com/IHTSDO/snomed-owl-toolkit * DeepOnto工具:https://github.com/KRR-Oxford/DeepOnto(基于OWLAPI(https://owlapi.sourceforge.net/),用于本体处理与复杂概念的自然语言表述)

提供机构:
Zenodo
创建时间:
2023-08-09
二维码
社区交流群
二维码
科研交流群
商业服务