OMOP results as of 20/10/22.
收藏资源简介:
BackgroundThe use of routinely collected health data for secondary research purposes is increasingly recognised as a methodology that advances medical research, improves patient outcomes, and guides policy. This secondary data, as found in electronic medical records (EMRs), can be optimised through conversion into a uniform data structure to enable analysis alongside other comparable health metric datasets. This can be achieved with the Observational Medical Outcomes Partnership Common Data Model (OMOP-CDM), which employs a standardised vocabulary to facilitate systematic analysis across various observational databases. The concept behind the OMOP-CDM is the conversion of data into a common format through the harmonisation of terminologies, vocabularies, and coding schemes within a unique repository. The OMOP model enhances research capacity through the development of shared analytic and prediction techniques; pharmacovigilance for the active surveillance of drug safety; and ‘validation’ analyses across multiple institutions across Australia, the United States, Europe, and the Asia Pacific. In this research, we aim to investigate the use of the open-source OMOP-CDM in the PATRON primary care data repository.MethodsWe used standard structured query language (SQL) to construct, extract, transform, and load scripts to convert the data to the OMOP-CDM. The process of mapping distinct free-text terms extracted from various EMRs presented a substantial challenge, as many terms could not be automatically matched to standard vocabularies through direct text comparison. This resulted in a number of terms that required manual assignment. To address this issue, we implemented a strategy where our clinical mappers were instructed to focus only on terms that appeared with sufficient frequency. We established a specific threshold value for each domain, ensuring that more than 95% of all records were linked to an approved vocabulary like SNOMED once appropriate mapping was completed. To assess the data quality of the resultant OMOP dataset we utilised the OHDSI Data Quality Dashboard (DQD) to evaluate the plausibility, conformity, and comprehensiveness of the data in the PATRON repository according to the Kahn framework.ResultsAcross three primary care EMR systems we converted data on 2.03 million active patients to version 5.4 of the OMOP common data model. The DQD assessment involved a total of 3,570 individual evaluations. Each evaluation compared the outcome against a predefined threshold. A ’FAIL’ occurred when the percentage of non-compliant rows exceeded the specified threshold value. In this assessment of the primary care OMOP database described here, we achieved an overall pass rate of 97%.ConclusionThe OMOP CDM’s widespread international use, support, and training provides a well-established pathway for data standardisation in collaborative research. Its compatibility allows the sharing of analysis packages across local and international research groups, which facilitates rapid and reproducible data comparisons. A suite of open-source tools, including the OHDSI Data Quality Dashboard (Version 1.4.1), supports the model. Its simplicity and standards-based approach facilitates adoption and integration into existing data processes.
背景 将常规收集的健康数据用于二次研究,现已日益被视为推动医学研究进步、改善患者预后并指导政策制定的研究方法。电子病历(Electronic Medical Records, EMRs)中存储的此类二次数据,可通过转换为统一数据结构得到优化,从而能够与其他可比健康指标数据集开展联合分析。这一目标可通过观察性医疗结果合作组织通用数据模型(Observational Medical Outcomes Partnership Common Data Model, OMOP-CDM)实现,该模型采用标准化词汇表,以支持跨多观察数据库的系统性分析。OMOP-CDM的核心理念是通过统一术语、词汇表与编码方案,将各类数据整合至单一数据仓库并转换为通用格式。该模型通过开发共享分析与预测技术、用于药物安全性主动监测的药物警戒(pharmacovigilance)流程,以及在澳大利亚、美国、欧洲与亚太地区多机构间开展的‘验证性’分析,提升了研究能力。本研究旨在探究开源OMOP-CDM在PATRON初级保健数据仓库中的应用。 方法 我们采用标准结构化查询语言(SQL)编写构建、提取、转换与加载脚本,将数据转换为OMOP-CDM格式。从各类电子病历中提取的自由文本术语的映射流程面临重大挑战:多数术语无法通过直接文本比对自动匹配至标准词汇表,导致大量术语需手动映射。为解决该问题,我们制定了一项策略,要求临床映射专员仅聚焦于出现频率足够高的术语。我们为每个数据领域设定了特定阈值,确保在完成合理映射后,超过95%的记录可匹配至SNOMED等经批准的标准词汇表。为评估生成的OMOP数据集的数据质量,我们采用OHDSI数据质量仪表板(OHDSI Data Quality Dashboard, DQD),依据Kahn框架对PATRON仓库中的数据的合理性、合规性与全面性进行评估。 结果 我们共将来自3个初级保健电子病历系统的203万活跃患者的数据转换为OMOP通用数据模型5.4版本。本次DQD评估共完成3570项独立评估,每项评估均将结果与预设阈值进行比对:当不合规记录行占比超过指定阈值时,即判定为‘不合格’。在本次针对上述初级保健OMOP数据库的评估中,我们获得了97%的整体通过率。 结论 OMOP-CDM在全球范围内的广泛应用、配套支持与培训,为协作研究中的数据标准化提供了成熟可行的路径。其兼容性支持本地与国际研究团队共享分析工具包,从而实现快速且可重复的数据比对。一系列开源工具,包括1.4.1版本的OHDSI数据质量仪表板,为该模型提供了支持。该模型凭借其简洁的设计与基于标准的实现思路,便于被现有数据流程采纳与集成。



