遇见数据集

N3C-Formatted OMOP2OBO Mappings

收藏
Zenodo2022-10-27 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>OMOP2OBO Mappings - N3C OMOP to OBO Working group</strong> This repository stores OMOP2OBO mappings which have been processed for use within the National COVID Cohort Collaborative (N3C) Enclave. The version of the mappings stored in this repository have been specifically formatted for use within the N3C Enclave. <strong>N3C OMOP to OBO Working Group: </strong>https://covid.cd2h.org/ontology <em><strong>Accessing the N3C-Formatted Mappings </strong></em> You can access the three OMOP2OBO HPO mapping files in the Enclave from the Knowledge store using the following link: https://unite.nih.gov/workspace/compass/view/ri.compass.main.folder.1719efcf-9a87-484f-9a67-be6a29598567. The mapping set includes three files, but you only need to merge the following two files with existing data in the Enclave in order to be able to create the concept sets: <em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv</em> <em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv</em> The first file <em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv</em>, contains columns for the OMOP concept ids and codes as well as specifies information like whether or not the OMOP concept’s descendants should be included when deriving the concept sets (defaults to FALSE). The other file <em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv</em>, contains details on the mapping’s label (i.e., the HPO curie and label in the concept_set_id field) and its provenance/evidence (the specific column to access for this information is called intention). <em><strong>Creating Concept Sets</strong></em> Merge these files together on the column named <em>codeset_id</em> and then join them with existing Enclave tables like concept and <em>condition_occurrence</em> to populate the actual concept sets. The name of the concept set can be obtained from the <em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv</em> file and is stored as a string in the column called <em>concept_set_id</em>. Although not ideal (but is the best way to approach this currently given what fields are available in the Enclave), to get the HPO CURIE and label will require applying a regex to this column. An example mapping is shown below (highlighting some of the most useful columns): <pre><code>codeset_id: 900000000 concept_set_id: [OMOP2OBO] hp_0002031-abnormal_esophagus_morphology concept: 23868 code: 69771008 codeSystem: SNOMED includeDescendants: False intention: Mixed - This mapping was created using the OMOP2OBO mapping algorithm (https://github.com/callahantiff/OMOP2OBO). The Mapping Category and Evidence supporting the mappings are provided below, by OMOP concept: 23868 ******* Mapping Category: Automatic Exact - Concept ------------------------------------------------ Mapping Provenance ------------------ OBO_DbXref-OMOP_ANCESTOR_SOURCE_CODE:snomed_69771008 | OBO_DbXref-OMOP_CONCEPT_SOURCE_CODE:snomed_69771008 | CONCEPT_SIMILARITY:HP_0002031_0.713</code></pre> <strong>Release Notes - v1.0.0</strong> <em>Preparation</em> In order to import data into the Enclave, the following items are needed: Obtain API Token, which will be included in the authorization header (stored as GitHub Secret) Obtain username hash from the Enclave OMOP2OBO Mappings (v1.0.0) <em>Data</em> Concept Set Container (<em>concept_set_container</em>): <em>CreateNewConceptSet</em> Concept Set Version (<em>code_sets</em>): C<em>reateNewDraftOMOPConceptSetVersion</em> Concept Set Expression Items (<em>concept_set_version_item</em>): <em>addCodeAsVersionExpression</em> <em>Script</em> <em>n3c_mapping_conversion.py</em>: https://github.com/callahantiff/OMOP2OBO/tree/master/applications/N3C <em>Generated Output</em> Need to have the <em>codeset_id</em> filled from self-generation (ideally, from a conserved range) prior to beginning any of the API steps. The current list of assigned identifiers is stored in the file named <em>omop2obo_enclave_codeset_id_dict_v1.0.0.json</em>. To be consistent with OMOP tools, specifically Atlas, we have also created Atlas-formatted json files for each mapping, which are stored in the zipped directory named <em>atlas_json_files_v1.0.0.zip</em>. <strong>File 1: concept_set_container</strong> <strong>Generated Data:</strong> <em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_container.csv</em> Columns: concept_set_id concept_set_name intention assigned_informatician assigned_sme project_id status stage n3c_reviewer alias archived created_by created_at <strong>File 2: concept_set_expression_items</strong> <strong>Generated Data: </strong><em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv</em> Columns: codeset_id concept_id code codeSystem isExcluded includeDescendants includeMapped item_id annotation created_by created_at <strong>File 3: concept_set_version</strong> <strong>Generated Data: </strong><em>OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv</em> Columns: codeset_id concept_set_id concept_set_version_title project source_application source_application_version created_at atlas_json most_recent_version comments intention limitations issues update_message status has_review reviewed_by created_by provenance atlas_json_resource_url parent_version_id is_draft <strong><em>Generated Output:</em></strong> OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_container.csv OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv atlas_json_files_v1.0.0.zip omop2obo_enclave_codeset_id_dict_v1.0.0.json

**OMOP2OBO映射集——N3C OMOP至OBO工作组** 本仓库存储经处理后可用于国家新冠队列协作组(National COVID Cohort Collaborative, N3C)安全分区(Enclave)的OMOP2OBO映射集。本仓库内存储的映射集版本已针对N3C安全分区的使用场景完成专属格式化。 **N3C OMOP至OBO工作组:**https://covid.cd2h.org/ontology *访问N3C适配版映射集* 您可通过以下链接,从N3C安全分区的知识库中获取三款OMOP2OBO人类表型本体(Human Phenotype Ontology, HPO)映射文件:https://unite.nih.gov/workspace/compass/view/ri.compass.main.folder.1719efcf-9a87-484f-9a67-be6a29598567。 该映射集共包含三个文件,但仅需将以下两个文件与N3C安全分区内的现有数据进行合并,即可构建概念集(concept sets): * `OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv` * `OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv` 第一个文件`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv`包含OMOP概念ID与编码相关字段,并标注了构建概念集时是否需纳入该OMOP概念的子概念(默认值为`FALSE`)。 另一文件`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv`则包含映射集标签(即`concept_set_id`字段中的HPO CURIE(统一资源标识符压缩格式,Compact Uniform Resource Identifier)与标签)及其溯源/佐证信息(可通过名为`intention`的字段获取该类信息)。 *概念集构建流程* 需以`codeset_id`列为关联键合并上述两个文件,随后将合并结果与N3C安全分区内的现有数据表(如`concept`表与`condition_occurrence`表)进行关联,以生成完整的概念集。 概念集名称可从`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv`文件中获取,该名称以字符串形式存储于`concept_set_id`字段中。尽管这并非最优方案,但鉴于当前N3C安全分区内可用字段的限制,这仍是目前获取HPO CURIE与标签的最佳方式——需对该字段应用正则表达式(regex)进行解析。以下为示例映射(展示了部分核心字段): codeset_id: 900000000 concept_set_id: [OMOP2OBO] hp_0002031-abnormal_esophagus_morphology concept: 23868 code: 69771008 codeSystem: SNOMED(系统医学术语集,Systematized Nomenclature of Medicine--Clinical Terms) includeDescendants: False intention: Mixed - This mapping was created using the OMOP2OBO mapping algorithm (https://github.com/callahantiff/OMOP2OBO). The Mapping Category and Evidence supporting the mappings are provided below, by OMOP concept: 23868 ******* Mapping Category: Automatic Exact - Concept ------------------------------------------------ Mapping Provenance ------------------ OBO_DbXref-OMOP_ANCESTOR_SOURCE_CODE:snomed_69771008 | OBO_DbXref-OMOP_CONCEPT_SOURCE_CODE:snomed_69771008 | CONCEPT_SIMILARITY:HP_0002031_0.713 **版本发布说明——v1.0.0** *准备工作* 若需将数据导入N3C安全分区,需完成以下准备事项: 1. 获取API令牌(API Token),该令牌将包含于授权请求头中,存储为GitHub密钥(GitHub Secret) 2. 从OMOP2OBO映射集(v1.0.0)的N3C安全分区文档中获取用户名哈希值 *数据文件* - 概念集容器(`concept_set_container`):`CreateNewConceptSet` - 概念集版本(`code_sets`):`CreateNewDraftOMOPConceptSetVersion` - 概念集表达式项(`concept_set_version_item`):`addCodeAsVersionExpression` *脚本文件* `n3c_mapping_conversion.py`:https://github.com/callahantiff/OMOP2OBO/tree/master/applications/N3C *生成的输出文件* 在执行任何API操作前,需先生成并填充`codeset_id`字段(理想情况下使用预留的标识符范围)。当前已分配的标识符列表存储于`omop2obo_enclave_codeset_id_dict_v1.0.0.json`文件中。为与OMOP工具(尤其是OMOP Atlas)保持兼容,我们还为每个映射生成了Atlas格式化的JSON文件,存储于压缩目录`atlas_json_files_v1.0.0.zip`内。 **文件1:concept_set_container** **生成数据:**`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_container.csv` 字段:`concept_set_id`、`concept_set_name`、`intention`、`assigned_informatician`、`assigned_sme`、`project_id`、`status`、`stage`、`n3c_reviewer`、`alias`、`archived`、`created_by`、`created_at` **文件2:concept_set_expression_items** **生成数据:**`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv` 字段:`codeset_id`、`concept_id`、`code`、`codeSystem`、`isExcluded`、`includeDescendants`、`includeMapped`、`item_id`、`annotation`、`created_by`、`created_at` **文件3:concept_set_version** **生成数据:**`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv` 字段:`codeset_id`、`concept_set_id`、`concept_set_version_title`、`project`、`source_application`、`source_application_version`、`created_at`、`atlas_json`、`most_recent_version`、`comments`、`intention`、`limitations`、`issues`、`update_message`、`status`、`has_review`、`reviewed_by`、`created_by`、`provenance`、`atlas_json_resource_url`、`parent_version_id`、`is_draft` ***生成的输出文件:*** `OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_container.csv`、`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_expression_items.csv`、`OMOP2OBO_v1.0.0_N3C_Enclave_CSV_concept_set_version.csv`、`atlas_json_files_v1.0.0.zip`、`omop2obo_enclave_codeset_id_dict_v1.0.0.json`

提供机构:
Zenodo
创建时间:
2022-10-25
二维码
社区交流群
二维码
科研交流群
商业服务