CRITICAL
收藏资源简介:
CRITICAL数据集是首个跨CTSA的倡议,创建了一个多站点、多模式、匿名化的临床数据集,结合了深度纵向覆盖和广泛的机构多样性。该数据集由四个CTSA站点(Northwestern、Tufts、华盛顿大学圣路易斯分校和阿拉巴马大学伯明翰分校)合作开发,包含了来自371,365名患者的19.5亿条记录,是迄今为止最大的公开共享、疾病无关的基准数据集,用于重症监护研究。CRITICAL数据集基于OMOP CDM v5.3,包含17张表,总计278.97GB,其中MEASUREMENT表包含14亿行。该数据集包括3800万次访问和2800万条单位级记录,平均每名患者有5,242行。CRITICAL提供了全面的病人护理历程,中位观察期为3.11年,最长跨度为31.8年,捕获了住院和门诊环境中ICU前、ICU和ICU后的遭遇。这种多机构、纵向的观点引入了大量的词汇异质性,需要在统一的标准下进行系统性的协调。CRITICAL数据集的这种大规模、机构多样性和纵向深度改善了AI研究人员对大规模临床数据的访问。
The CRITICAL dataset is the first cross-CTSA initiative that creates a multi-site, multi-modal, anonymized clinical dataset combining deep longitudinal coverage and broad institutional diversity. Developed collaboratively by four CTSA sites including Northwestern, Tufts, Washington University in St. Louis, and University of Alabama at Birmingham, the dataset contains 1.95 billion records from 371,365 patients, making it the largest publicly shared, disease-agnostic benchmark dataset for critical care research to date. Based on OMOP CDM v5.3, the dataset comprises 17 tables with a total size of 278.97 GB, among which the MEASUREMENT table contains 1.4 billion rows. It also covers 38 million visits and 28 million unit-level records, with an average of 5,242 rows per patient. The CRITICAL dataset provides a comprehensive view of patient care trajectories, with a median observation period of 3.11 years and a maximum span of 31.8 years, capturing pre-ICU, ICU, and post-ICU encounters in both inpatient and outpatient settings. This multi-institutional, longitudinal perspective introduces significant lexical heterogeneity, requiring systematic harmonization under a unified standard. The scale, institutional diversity and longitudinal depth of the CRITICAL dataset improve access to large-scale clinical data for AI researchers.
CRITICAL 数据集概述
数据集名称
CRITICAL(Collaborative Resource for Intensive-care Translational science, Informatics, Comprehensive Analytics, and Learning)
资助信息
- 资助机构:NIH National Center for Advancing Translational Sciences
- 奖项编号:U01TR003528
数据集目标
- 创建首个跨CTSA的多中心、多模态、去标识化数据集,兼具深度和广度。
- 解决临床人工智能(AI)转化研究中共享数据资源不足的问题。
数据内容
- 数据类型:纵向住院和门诊数据,包括ICU入院前和后的数据。
- 数据规模:来自超过40万名重症患者的临床数据。
- 数据特点:
- 目前最大的公开共享、疾病无关的基准临床数据集。
- 涵盖多样化的种族、民族和地理特征。
应用领域
- 人工智能(AI)和机器学习(ML)研究。
- 结局相关研究。
- 支持公平和可泛化的AI转化,用于高级患者监测和决策支持。
访问条件
- 适用对象:
- 美国认证大学的教师、学生和工作人员(个人申请)。
- 美国认证大学(机构申请)。
- 要求:机构需与CRITICAL Consortium签订数据使用协议(DUA)。
相关资源
- 数据访问:https://critical.fsm.northwestern.edu
- 数据代码书
- 常见问题解答(FAQ)
- 联系方式
参与机构
- 西北大学(Northwestern University)
- 西北医学(Northwestern Medicine)
- 其他CTSA站点:Tufts、WUSTL、UAB




