遇见数据集

Graph Neural Network and Sentence Transformer Embeddings for SNOMED CT concepts

收藏
Zenodo2026-02-02 更新2026-05-29 收录
官方服务:

资源简介:

Embeddings for SNOMED CT concepts produced by Graph Neural Networks (GNNs) or Sentence Transformer model. Each NPZ file encodes a dictionary, which links the ID of a SNOMED CT concept to its corresponding embedding. Files base_mini_lm_dict.npz and fine_tuned_mini_lm_dict.npz contain the embeddings of the sentence transformer models, where the former is using the base MiniLM model and the latter is using the fine-tuned MiniLM model on the concept similarity task. Files gnn_mul_sct_dict.npz and gnn_sim_sct_dict.npz contain the embeddings produced by a GNN on a dataset produced by transforming the SNOMED CT ontology and on the task of concept similarity. These embeddings were generated and studied in the paper Assessing the Effectiveness of Embedding Methods in Capturing Clinical Information from SNOMED CT () and more information can also be found in the following repository: https://github.com/JavierCastellD/AssessingSNOMEDEmbeddings.

本数据集包含由图神经网络(Graph Neural Networks,GNNs)或句子Transformer(Sentence Transformer)模型生成的医学系统命名法临床术语集(SNOMED CT)概念嵌入向量。每个NPZ文件均编码一个字典,该字典将SNOMED CT概念的唯一标识符与其对应的嵌入向量相关联。文件base_mini_lm_dict.npz与fine_tuned_mini_lm_dict.npz存储了句子Transformer模型生成的嵌入向量:其中前者基于基础版MiniLM模型,后者则是在概念相似度任务上微调后的MiniLM模型所生成的嵌入向量。文件gnn_mul_sct_dict.npz与gnn_sim_sct_dict.npz则存储了图神经网络生成的嵌入向量,这些向量是基于经过转换的SNOMED CT本体数据集,并针对概念相似度任务生成的。上述嵌入向量已在论文《Assessing the Effectiveness of Embedding Methods in Capturing Clinical Information from SNOMED CT》中完成生成与研究,更多相关信息可通过以下开源仓库获取:https://github.com/JavierCastellD/AssessingSNOMEDEmbeddings。

提供机构:
Zenodo
创建时间:
2026-02-02
二维码
社区交流群
二维码
科研交流群
商业服务