遇见数据集

A2H_Clinical_Data

收藏
Zenodo2026-03-31 更新2026-05-26 收录
官方服务:

资源简介:

General The two files in this ressource are used in the analysis of animal to human translation for this project: Preclinical_DrugDisease_Translation_Pipeline. Raw AACT Snapshot raw_aact/mv_interventional_drug_studies_20260302.csvTabular snapshot of interventional drug-related studies derived from the AACT / ClinicalTrials.gov relational database. The file was generated from a materialized view built on a database snapshot dated 1 December 2025. It includes one row per nct_id for studies with study_type = 'INTERVENTIONAL' and at least one intervention of type DRUG, DIETARY_SUPPLEMENT, BIOLOGICAL, COMBINATION_PRODUCT, GENETIC, or OTHER. The table combines study-level metadata from ctgov.studies, brief summaries from ctgov.brief_summaries, aggregated intervention names and types from ctgov.interventions, and aggregated condition names from ctgov.conditions. Included columns: nct_id brief_title study_official_title start_date completion_date study_first_submitted_date phase overall_status brief_summary intervention_names intervention_types condition_names Notes: intervention_names, intervention_types, and condition_names are aggregated as pipe-separated strings (" | "). Only studies matching the SQL selection criteria are included. This file is intended as the structured trial metadata input for downstream entity-linking and integration steps. Linked NER Drug and Disease Entities linked_to_ontologies/entities_drug_disease_clin.csvNormalized drug and disease ontology annotations applied to the NER results. Disease concepts are mapped to MONDO, while drug concepts are mapped to UMLS CUIs. Multiple entities are represented as pipe-separated values (|). Disease / condition mapping (MONDO) merged_condition_names: Original condition names aggregated from the trial record disease_mondo_termid: Assigned MONDO identifier disease_mondo_term_norm: Normalized MONDO label disease_term_mondo_clean: Cleaned disease string used for matching disease_termid_mondo_clean: MONDO ID after cleaning step nearest_dataset_parent_mondo: Closest parent MONDO concept in the reference dataset (-1 if none) nearest_dataset_parent_label: Label of the nearest parent concept merged_mondo_termid: Final merged MONDO identifier(s) merged_mondo_label: Final merged MONDO label(s) Drug / intervention mapping (UMLS) ner_predicted_drugs: Drug names extracted via NER linkbert_umls_drugs: Drug names after normalization / linking model drug_umls_termid: UMLS concept identifiers (CUIs) drug_umls_term_norm: Normalized UMLS labels nearest_dataset_parent_umls: Closest parent UMLS concept (-1 if none) nearest_dataset_parent_umls_label: Label of the parent concept merged_umls_termid: Final merged UMLS identifier(s) merged_umls_label: Final merged UMLS label(s)

# 概述 本资源包含的两个文件,用于本项目**Preclinical_DrugDisease_Translation_Pipeline**(临床前药物-疾病翻译管线)的动物向人类翻译分析工作。 ## 原始AACT快照 `raw_aact/mv_interventional_drug_studies_20260302.csv` 该文件为源自AACT/ClinicalTrials.gov关系型数据库的干预性药物相关研究的结构化快照,基于2025年12月1日的数据库快照构建的实体化视图生成。其每一行对应一个`nct_id`,筛选条件为研究类型为`INTERVENTIONAL`(干预性研究)且至少包含一项类型为`DRUG`(药物)、`DIETARY_SUPPLEMENT`(膳食补充剂)、`BIOLOGICAL`(生物制品)、`COMBINATION_PRODUCT`(联合产品)、`GENETIC`(基因治疗产品)或`OTHER`(其他)的干预措施。该表格整合了`ctgov.studies`中的研究级元数据、`ctgov.brief_summaries`中的简要摘要、`ctgov.interventions`中的干预名称与类型聚合信息,以及`ctgov.conditions`中的疾病名称聚合信息。 ### 包含字段: `nct_id`、`brief_title`(研究简要标题)、`study_official_title`(研究正式标题)、`start_date`(启动日期)、`completion_date`(完成日期)、`study_first_submitted_date`(首次提交研究日期)、`phase`(研究阶段)、`overall_status`(整体研究状态)、`brief_summary`(简要摘要)、`intervention_names`(干预名称)、`intervention_types`(干预类型)、`condition_names`(疾病名称) ### 说明: 1. `intervention_names`、`intervention_types`与`condition_names`以竖线(`|`)作为分隔符进行聚合存储。 2. 仅包含符合SQL筛选条件的研究数据。 3. 该文件旨在作为下游实体链接与集成任务的结构化试验元数据输入文件。 ## 关联命名实体识别(Named Entity Recognition, NER)药物与疾病实体 `linked_to_ontologies/entities_drug_disease_clin.csv` 该文件为应用于命名实体识别结果的标准化药物与疾病本体注释。其中疾病概念映射至MONDO本体,药物概念映射至UMLS的CUI(概念唯一标识符)。多个实体以竖线(`|`)分隔的字符串形式存储。 ### 疾病/条件映射(MONDO) - `merged_condition_names`: 从试验记录中聚合得到的原始疾病名称 - `disease_mondo_termid`: 分配的MONDO标识符 - `disease_mondo_term_norm`: 标准化后的MONDO标签 - `disease_term_mondo_clean`: 用于匹配的清洗后疾病字符串 - `disease_termid_mondo_clean`: 清洗步骤后的MONDO ID - `nearest_dataset_parent_mondo`: 参考数据集中最接近的父级MONDO概念(若无则为`-1`) - `nearest_dataset_parent_label`: 最接近的父级概念的标签 - `merged_mondo_termid`: 最终合并的MONDO标识符 - `merged_mondo_label`: 最终合并的MONDO标签 ### 药物/干预映射(UMLS) - `ner_predicted_drugs`: 通过命名实体识别提取的药物名称 - `linkbert_umls_drugs`: 经过标准化/链接模型处理后的药物名称 - `drug_umls_termid`: UMLS概念标识符(CUIs) - `drug_umls_term_norm`: 标准化后的UMLS标签 - `nearest_dataset_parent_umls`: 参考数据集中最接近的父级UMLS概念(若无则为`-1`) - `nearest_dataset_parent_umls_label`: 父级概念的标签 - `merged_umls_termid`: 最终合并的UMLS标识符 - `merged_umls_label`: 最终合并的UMLS标签

提供机构:
Zenodo
创建时间:
2026-03-31
二维码
社区交流群
二维码
科研交流群
商业服务