遇见数据集

DataTecnica/RoP_biomedical

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

RoP(生物医学参考参数集)是一个包含133万个预协调通用数据元素(CDEs)的大型数据集,涵盖生物医学数据的13个主题,包括临床表型、基因组学、影像学、生物样本、评估工具、治理和资源目录等。它整合了OMOP、LOINC、ICD-10、RxNorm、HPO、Mondo、NINDS-CDE、PhenX、CDISC、DICOM、BIDS、DUO等多个主要生物医学词汇表,通过AI驱动的语义嵌入(如SapBERT)实现快速CDE匹配,支持数据互操作性和AI就绪性。数据集采用FAIR设计原则,版本化季度更新,并已在实际生产环境中应用于大规模联合开放科学项目。其技术深度包括139万个通用数据模型组件协调为133万个互操作数据元素,映射到代表4000多个语义集群的生物医学语言模型嵌入。数据集旨在解决多队列研究中数据协调的瓶颈,减少人工映射时间,促进生物医学研究的加速发展。

RoP (Biomedical Reference Parameter Set) is a large-scale dataset containing 1.33 million pre-coordinated common data elements (CDEs), covering 13 thematic domains of biomedical data including clinical phenotypes, genomics, imaging, biospecimens, assessment tools, governance, resource catalogs, and more. It integrates multiple major biomedical vocabularies such as OMOP, LOINC, ICD-10, RxNorm, HPO, Mondo, NINDS-CDE, PhenX, CDISC, DICOM, BIDS, and DUO, and enables fast CDE matching via AI-driven semantic embeddings (e.g., SapBERT) to support data interoperability and AI readiness. The dataset adheres to the FAIR guiding principles, is versioned and updated quarterly, and has been deployed in real-world production environments for large-scale collaborative open science projects. Its technical depth lies in the coordination of 1.39 million common data model components into 1.33 million interoperable data elements, which are mapped to biomedical language model embeddings representing over 4,000 semantic clusters. The dataset aims to resolve data harmonization bottlenecks in multi-cohort studies, reduce manual mapping time, and promote the accelerated advancement of biomedical research.

提供机构:
DataTecnica
二维码
社区交流群
二维码
科研交流群
商业服务