遇见数据集

AI-derived and Manually corrected segmentations for various IDC Collections

收藏
Zenodo2024-09-27 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

The Imaging Data Commons (IDC)(https://imaging.datacommons.cancer.gov/) [1] connects researchers with publicly available cancer imaging data, often linked with other types of cancer data. Many of the collections have limited annotations due to the expense and effort required to create these manually. The increased capabilities of AI analysis of radiology images provides an opportunity to augment existing IDC collections with new annotation data. To further this goal, we trained several nnUNet [2] based models for a variety of radiology segmentation tasks from public datasets and used them to generate segmentations for IDC collections. To validate the models performance, roughly 10% of the predictions were manually reviewed and corrected by both a board certified radiologist and a medical student (non-expert). Additionally, this non-expert looked at all the ai predictions and rated them on a 5 point Likert scale . This record provides the AI segmentations, Manually corrected segmentations, and Manual scores for the inspected IDC Collection images. <strong>List of all tasks and IDC collections analyzed.</strong> File Segmentation Task IDC Collections LInks breast-fdg-pet-ct.zip FDG-avid lesions in breast from FDG PET/CT scans QIN-Breast model weights github kidney-ct.zip Kidney, Tumor, and Cysts from contrast enhanced CT scans TCGA-KIRC model weights github liver-ct.zip Liver from CT scans TCGA-LIHC model weights github liver-mr.zip Liver from T1 MRI scans TCGA-LIHC model weights github lung-ct.zip Lung and Nodules (3mm-30mm) from CT scans ACRIN-NSCLC-FDG-PET<br> Anti-PD-1-Lung<br> LUNG-PET-CT-Dx<br> NSCLC Radiogenomics<br> RIDER Lung PET-CT<br> TCGA-LUAD<br> TCGA-LUSC model weights 1 model weights 2 github lung-fdg-pet-ct.zip Lungs and FDG-avid lesions in the lung from FDG PET/CT scans ACRIN-NSCLC-FDG-PET<br> Anti-PD-1-Lung<br> LUNG-PET-CT-Dx<br> NSCLC Radiogenomics<br> RIDER Lung PET-CT<br> TCGA-LUAD<br> TCGA-LUSC model weights github prostate-mr.zip Prostate from T2 MRI scans ProstateX model weights github Likert Score Definition 5 Strongly Agree - Use-as-is (i.e., clinically acceptable, and could be used for treatment without change) 4 Agree - Minor edits that are not necessary. Stylistic differences, but not clinically important. The current segmentation is acceptable 3 Neither agree nor disagree - Minor edits that are necessary. Minor edits are those that the review judges can be made in less time than starting from scratch or are expected to have minimal effect on treatment outcome 2 Disagree - Major edits. This category indicates that the necessary edit is required to ensure correctness, and sufficiently significant that user would prefer to start from the scratch 1 Strongly disagree - Unusable. This category indicates that the quality of the automatic annotations is so bad that they are unusable. Each zip file in the collection correlates to a specific segmentation task. The common folder structure is ai-segmentations-dcm This directory contains the AI model predictions in DICOM-SEG format for all analyzed IDC collection files qa-segmentations-dcm This directory contains manual corrected segmentation files, based on the AI prediction, in DICOM-SEG format. Only a fraction, ~10%, of the AI predictions were corrected. Corrections were performed by radiologist (rad*) and non-experts (ne*) qa-results.csv CSV file linking the study/series UIDs with the ai segmentation file, radiologist corrected segmentation file, radiologist ratings of AI performance.

成像数据共同体(Imaging Data Commons, IDC)[1](https://imaging.datacommons.cancer.gov/)将研究人员与公开可用的癌症影像数据进行对接,此类数据通常还关联了其他类型的癌症组学数据。由于手动标注需耗费高额成本与大量人力,多数数据集集合仅带有有限的标注信息。当前AI影像分析能力的提升为利用新增标注数据扩充现有IDC数据集集合提供了全新契机。为推进这一目标,我们基于nnUNet[2]训练了多款模型,针对来自公开数据集的多种放射学分割任务开展训练,并使用这些模型为IDC数据集集合生成分割结果。为验证模型性能,我们由一名执业放射科医师与一名医学生(非专业标注者)对约10%的预测结果进行了人工审阅与修正。此外,该非专业标注者还对所有AI预测结果进行了评分,评分采用5级李克特(Likert)量表。本数据集包含AI生成分割结果、人工修正后的分割结果以及经审阅的IDC集合图像的人工评分。<strong>所有分析任务与IDC数据集集合列表如下。</strong> | 分割任务压缩包 | 任务描述 | 关联IDC数据集集合 | 模型权重链接 | | --- | --- | --- | --- | | breast-fdg-pet-ct.zip | 针对FDG PET/CT扫描中乳腺FDG高摄取病灶的分割 | QIN-Breast | GitHub | | kidney-ct.zip | 增强CT扫描中肾脏、肿瘤与囊肿的分割 | TCGA-KIRC | GitHub | | liver-ct.zip | CT扫描中肝脏的分割 | TCGA-LIHC | GitHub | | liver-mr.zip | T1 MRI扫描中肝脏的分割 | TCGA-LIHC | GitHub | | lung-ct.zip | CT扫描中肺部与(3mm-30mm)结节的分割 | ACRIN-NSCLC-FDG-PET、Anti-PD-1-Lung、LUNG-PET-CT-Dx、NSCLC Radiogenomics、RIDER Lung PET-CT、TCGA-LUAD、TCGA-LUSC | 权重1、权重2、GitHub | | lung-fdg-pet-ct.zip | 针对FDG PET/CT扫描中肺部与肺内FDG高摄取病灶的分割 | ACRIN-NSCLC-FDG-PET、Anti-PD-1-Lung、LUNG-PET-CT-Dx、NSCLC Radiogenomics、RIDER Lung PET-CT、TCGA-LUAD、TCGA-LUSC | GitHub | | prostate-mr.zip | T2 MRI扫描中前列腺的分割 | ProstateX | GitHub | ### 李克特量表评分定义 5分:完全同意——可直接使用(即临床可接受,无需修改即可用于临床治疗决策) 4分:同意——仅需进行非必要的小幅编辑。仅存在风格差异,不影响临床诊疗意义,当前分割结果可接受 3分:中立——需进行必要的小幅编辑。此类编辑指审阅者判断可在短于从头标注的时间内完成,且对治疗结局影响极小的修改 2分:不同意——需进行大幅编辑。该类别表示需进行必要修改以确保结果正确性,且修改幅度大到使用者更倾向于从头开始标注 1分:完全不同意——无法使用。该类别表示自动标注结果质量极差,完全不可用 ### 数据集目录结构 本数据集的每个压缩包均对应一项特定的分割任务。通用目录结构如下: 1. `ai-segmentations-dcm`:该目录包含所有待分析IDC集合文件对应的DICOM-SEG格式AI模型预测分割结果 2. `qa-segmentations-dcm`:该目录包含基于AI预测结果经人工修正后的DICOM-SEG格式分割文件。仅约10%的AI预测结果被修正,修正工作由放射科医师(rad*)与非专业标注者(ne*)完成 3. `qa-results.csv`:该CSV文件将研究/序列唯一标识符(UID)与AI分割文件、放射科医师修正后的分割文件、放射科医师对AI性能的评分相关联。

提供机构:
Zenodo
创建时间:
2023-09-16
二维码
社区交流群
二维码
科研交流群
商业服务