UngLong/openm3chest-labels-v2
收藏资源简介:
OpenM3Chest Labels v2是一个用于医疗人工智能的数据集,专门设计用于对MedGemma 1.5-4B-IT模型进行LoRA微调。该数据集基于OpenM3Chest项目,源自美国国家肺癌筛查试验(NLST),包含胸部CT扫描图像的标签子集,并覆盖多个医疗任务领域:放射学(如胸部异常检测、结节存在、位置、衰减、边缘和大小分类)、心脏病学(如心血管疾病诊断和死亡率预测)以及肿瘤学(如肺癌风险预测)。数据集提供分层采样,训练集针对二进制任务保持约50/50的平衡,并覆盖所有分类类别;测试集则使用完整的原始测试集以进行准确评估。每个配置(如chest_abn_54、CVD_diagnosis、lung_cancer_risk)对应一个特定任务,包含训练和测试分割,记录字段包括键(SeriesInstanceUID)、患者ID、临床数据(JSON格式)、问题列表、标签(整数或字符串)和答案字典(JSON格式)。CT扫描图像以NPY格式存储在另一个仓库中,可通过患者ID和键映射加载。数据集支持顺序推理,特别是在放射学任务中,首先进行筛查,然后根据结节存在结果进行进一步分析。数据来源包括OpenM3Chest和IDC(成像数据共享平台),旨在促进医疗影像和临床数据的多任务学习研究。
OpenM3Chest Labels v2 is a dataset for medical AI, specifically designed for LoRA fine-tuning of the MedGemma 1.5-4B-IT model. It is based on the OpenM3Chest project, derived from the National Lung Screening Trial (NLST), and contains a labeled subset of chest CT scan images covering multiple medical task domains: radiology (e.g., chest abnormality detection, nodule presence, location, attenuation, margin, and size classification), cardiology (e.g., CVD diagnosis and mortality prediction), and oncology (e.g., lung cancer risk prediction). The dataset provides stratified sampling, with the training set maintaining a ~50/50 balance for binary tasks and covering all categorical classes; the test set uses the full original test set for proper evaluation. Each configuration (e.g., chest_abn_54, CVD_diagnosis, lung_cancer_risk) corresponds to a specific task and includes train and test splits, with record fields such as keys (SeriesInstanceUID), patient IDs, clinical data (JSON-serialized), a list of questions, labels (integer or string), and an answer dictionary (JSON-serialized). CT scan images are stored in NPY format in a separate repository and can be loaded via patient ID and key mapping. The dataset supports sequential inference, particularly for radiology tasks, where screening is performed first, followed by further analysis based on nodule presence results. Data sources include OpenM3Chest and the Imaging Data Commons (IDC), aiming to facilitate multi-task learning research in medical imaging and clinical data.




