遇见数据集

Execution time of ETL process for example data 3.

收藏
Figshare2025-01-06 更新2026-04-28 收录
官方服务:

资源简介:

ObjectiveThe German Health Data Lab is going to provide access to German statutory health insurance claims data ranging from 2009 to the present for research purposes. Due to evolving data formats within the German Health Data Lab, there is a need to standardize this data into a Common Data Model to facilitate collaborative health research and minimize the need for researchers to adapt to multiple data formats. For this purpose we selected transforming the data to the Observational Medical Outcomes Partnership Common Data Model.MethodsWe developed an Extract, Transform, and Load (ETL) pipeline for two distinct German Health Data Lab data formats: Format 1 (2009-2016) and Format 3 (2019 onwards). Due to the identical format structure of Format 1 and Format 2 (2017 -2018), the ETL pipeline of Format 1 can be applied on Format 2 as well. Our ETL process, supported by Observational Health Data Sciences and Informatics tools, includes specification development, SQL skeleton creation, and concept mapping. We detail the process characteristics and present a quality assessment that includes field coverage and concept mapping accuracy using example data.ResultsFor Format 1, we achieved a field coverage of 92.7%. The Data Quality Dashboard showed 100.0% conformance and 80.6% completeness, although plausibility checks were disabled. The mapping coverage for the Condition domain was low at 18.3% due to invalid codes and missing mappings in the provided example data. For Format 3, the field coverage was 86.2%, with Data Quality Dashboard reporting 99.3% conformance and 75.9% completeness. The Procedure domain had very low mapping coverage (2.2%) due to the use of mocked data and unmapped local concepts The Condition domain results with 99.8% of unique codes mapped. The absence of real data limits the comprehensive assessment of quality.ConclusionThe ETL process effectively transforms the data with high field coverage and conformance. It simplifies data utilization for German Health Data Lab users and enhances the use of OHDSI analysis tools. This initiative represents a significant step towards facilitating cross-border research in Europe by providing publicly available, standardized ETL processes (https://github.com/FraunhoferMEVIS/ETLfromHDLtoOMOP) and evaluations of their performance.

研究目标:德国健康数据实验室(German Health Data Lab)将为科研用途提供2009年至今的德国法定健康保险索赔数据。鉴于德国健康数据实验室内部数据格式持续迭代更新,亟需将此类数据标准化为通用数据模型,以推动跨机构健康协作研究,并降低研究人员适配多种数据格式的工作量。为此,我们选择将数据转换为观察性医疗结果合作伙伴通用数据模型(Observational Medical Outcomes Partnership Common Data Model,简称OMOP CDM)。 研究方法:我们针对两类差异化的德国健康数据实验室数据格式开发了抽取-转换-加载(Extract, Transform, and Load,简称ETL)流水线:格式1(2009年至2016年)与格式3(2019年至今)。由于格式1与格式2(2017年至2018年)的结构完全一致,格式1的ETL流水线同样可适配格式2的数据。本ETL流程依托观察性健康数据科学与信息学(Observational Health Data Sciences and Informatics,简称OHDSI)工具开发,涵盖规范制定、SQL骨架构建与概念映射三大环节。本文详细阐述了该流程的核心特性,并基于示例数据呈现了包含字段覆盖率与概念映射准确率的质量评估结果。 研究结果:针对格式1,其字段覆盖率达92.7%。数据质量仪表盘(Data Quality Dashboard)检测结果显示,尽管未启用合理性检查环节,但数据符合率为100.0%,数据完整率为80.6%。由于示例数据中存在无效编码与未映射条目,疾病领域(Condition domain)的映射覆盖率仅为18.3%,处于较低水平。针对格式3,其字段覆盖率为86.2%,数据质量仪表盘报告显示符合率为99.3%,完整率为75.9%。由于使用了模拟数据且存在未映射的本地概念,操作领域(Procedure domain)的映射覆盖率极低,仅为2.2%;而疾病领域的映射覆盖率达99.8%的唯一编码均已完成映射。由于缺乏真实数据,无法开展全面的质量评估。 研究结论:本ETL流程可高效完成数据转换,具备较高的字段覆盖率与数据符合率,能够简化德国健康数据实验室用户的数据使用流程,并提升OHDSI分析工具的应用效能。本项目通过公开可获取的标准化ETL流程(https://github.com/FraunhoferMEVIS/ETLfromHDLtoOMOP)及其性能评估结果,为推动欧洲跨境健康研究迈出了重要一步。

创建时间:
2025-01-06
二维码
社区交流群
二维码
科研交流群
商业服务