融合国际标准的医学语言系统的知识库
收藏资源简介:
本数据集面向卫生健康领域医学语言体系不统一、术语标准差异导致的数据互通与共享障碍,围绕医学语言标准化与语义一致性需求,构建了一个大规模、跨标准的医学语言资源体系。本数据集文件均为CSV格式,由3个文件组成。其中data.csv文件为知识图谱三元组信息,采用整数编号形式对实体与关系进行统一编码。entity_dict.csv和relation_dict.csv分别为实体与关系的数据词典,提供实数据索引id与数据内容的对照。数据内容以三元组结构存储,用于表示知识图谱中医学概念术语的关联关系,数据总量包含12,853,046例医学概念和188,848,003条三元组记录,文件合计大小为13.3G。
This dataset addresses the barriers to data interconnection and sharing caused by inconsistent medical language systems and disparate terminology standards in the healthcare domain. It constructs a large-scale, cross-standard medical language resource system targeting the requirements of medical language standardization and semantic consistency. This dataset includes three files, all in CSV format. Among them, data.csv stores knowledge graph triple information, and uniformly encodes entities and relations using integer numbering. The entity_dict.csv and relation_dict.csv serve as the data dictionaries for entities and relations respectively, providing the mapping between their index IDs and their actual content. The data content is stored in triple structures, which are used to represent the associative relationships between medical conceptual terms in the knowledge graph. The total dataset contains 12,853,046 medical concepts and 188,848,003 triple records, with a combined file size of 13.3 GB.




