遇见数据集

CURVAS dataset

收藏
Zenodo2024-09-16 更新2026-05-26 收录
官方服务:

资源简介:

Clinical Problem In medical imaging, DL models are often tasked with delineating structures or abnormalities within complex anatomical structures, such as tumors, blood vessels, or organs. Uncertainty arises from the inherent complexity and variability of these structures, leading to challenges in precisely defining their boundaries. This uncertainty is further compounded by interrater variability, as different medical experts may have varying opinions on where the true boundaries lie. DL models must grapple with these discrepancies, leading to inconsistencies in segmentation results across different annotators and potentially impacting diagnosis and treatment decisions. Addressing interrater variability in DL for medical segmentation involves the development of robust algorithms capable of capturing and quantifying uncertainty, as well as standardizing annotation practices and promoting collaboration among medical experts to reduce variability and improve the reliability of DL-based medical image analysis. Interrater variability poses significant challenges in the field of DL for medical image segmentation. Furthermore, achieving model calibration, a fundamental aspect of reliable predictions, becomes notably challenging when dealing with multiple classes and raters. Calibration is pivotal for ensuring that predicted probabilities align with the true likelihood of events, enhancing the model's reliability. It must be considered that, even if not clearly, having multiple classes account for uncertainties arising from their interactions. Moreover, incorporating annotations from multiple raters adds another layer of complexity, as differing expert opinions may contribute to a broader spectrum of variability and computational complexity. Consequently, the development of robust algorithms capable of effectively capturing and quantifying variability and uncertainty, while also accommodating the nuances of multi-class and multi-rater scenarios, becomes imperative. Striking a balance between model calibration, accurate segmentation and handling variability in medical annotations is crucial for the success and reliability of DL-based medical image analysis. CURVAS Challenge Goal Due to all the previously stated reasons, we have created a challenge that considers all of the above. In this challenge, we will work with abdominal CT scans. Each of them will have three different annotations obtained from different experts and each of the annotations will have three classes: pancreas, kidney and liver. The main idea is to be able to evaluate the results considering the multi rater information. There will be three separate evaluations: firstly, a classical dice score evaluation together with an uncertainty study will be performed; secondly, a volumetric assessment to give relevant clinical information will take place; finally, a study on whether the model is calibrated or not will take place. All of these evaluations will be performed considering all three different annotations. For more information about the challenge, visit our website to join CURVAS (Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation). This challenge will be held in MICCAI 2024. Dataset Cohort The challenge cohort consists of 90 CT images prospectively gathered at the University Hospital Erlangen between August 2023 and October 2023. Each CT will have multiple classes: background (0), pancreas (1), kidney (2) and liver (3). In addition, each of the CTs will have three different annotators from three different experts that will contain the four classes specified previously. Training Phase cohort: 20 CT scans belonging to group A with the respective annotations will be given. It is encouraged to leverage publicly available external data annotated by multiple raters. The idea of giving a small amount of data for the training set and giving the opportunity of using a public dataset for training is to make the challenge more inclusive, giving the option to develop a method by using data that is in anyone's hands. Furthermore, by using this data to train and using other data to evaluate, it makes it more robust to shifts and other sources of variability between datasets. Validation Phase cohort: 5 CT scans belonging to group A will be used for this phase. Test Phase cohort: 65 CT scans will be used for evaluation. 20 CTs belonging to group A, 22 CTs belonging to group B and 23 CTs belonging to group C. Both validation and testing CT scans cohorts will not be published until the end of the challenge. Furthermore, to which group each CT scan belongs will not be revealed until after the challenge. Clinical Specifications Inclusion criteria were a maximum of 10 cysts with a diameter of less than 2,0 cm. Furthermore, CT scans with major artifacts (e.g. breathing artifacts) or incomplete registrations were excluded. Participants were required to be over 18 years old and provide both verbal and written consent for the use of their CT images in the Challenge. Both study-specific and broad consent were obtained. Among the 90 patients, there were 51 males and 39 females, aged between 37 and 94 years, with an average age of 65.7 years. All patients received treatment at the University Hospital Erlangen in Bavaria, Germany. No additional selection criteria was set to ensure a representative sample of a typical patient cohort. Our overall data consists on 90 CTs splitted in three different groups: Group A: cases with 2 cysts or less with no contour altering pathologies - 45 CTs Group B: cases with 3-5 cysts with no contour altering pathologies - 22 CTs Group C: cases with 6-10 cysts with some pathologies included (liver metastases, hydronephrosis, adrenal gland metastases, missing kidney) - 23 CTs However, in any case, the participants will not know which case belongs to which group. This information will be released after the challenge, together with the whole dataset. Annotation Protocol The first step for obtaining de labels was using the TotalSegmentator [1] [2] to get rough annotations. Then, the labels were sent to three radiologists (R1, R2, R3), to both correct the automatic annotations and add possible missing organs. One of the three labeling radiologists, the MD PhD candidate, previously defined both the dataset cohort and the criteria of what belongs to the parenchyma and what does not and it was given to the other two labeling radiologists to follow the same criteria to be coherent with each other [3]. Separately, two other clinicians (C1, C2) supervised the criteria of the cohort defined by the MD PhD candidate, but not having any relation with the labeling itself, hence, there is no bias between the annotations of the different radiologists. Each labeled class for this challenge has specific instructions. Below are listed per organ. Liver:Generally speaking, we define the liver 'as the entire liver tissue including all internal structures like vessel systems, tumors etc.' [4] Thus, the portal vein itself is excluded from contouring. The two main branches of the portal vein are excluded from the segmentation. Any branch of the following generations is included. 'In case of partial enclosure (occurring where large vessels as Vena Cava and portal vein enter or leave the liver), the parts enclosed by liver tissue are included in the segmentation, thus forming the convex hull of the liver shape.' [4] Any fatty tissue that pulls into the liver is excluded. The gallbladder should not be marked. Wide and especially pathologically widened bile ducts are included in the segmentation of the liver. Kidney:The right and left kidney will be segmented. Included in the segmentation will be the kidney parenchyma including the renal medulla. Excluded is the renal pelvis [5] and the ureter as a urinary stasis could alter the original volume. Pancreas:When segmenting the pancreas, we will not differentiate between head, body and tail. Moreover neither the splenic vein nor the mesenterial vein will be included in segmentation [6]. However, it is important the whole pancreas in its course is tracked and marked. Technical Specifications The CTs used needed to be contrast-enhanced CT scans in a portal venous phase with the acquisition of thin slices ranging from 0.6 to 1mm. Thoracic-Abdominal CT images were taken during the patients' hospital stay, motivated by various medical needs. Given the focus on abdominal organs, the Br40 soft kernel was employed. CT examinations were conducted using SIEMENS CT scanners at the university hospital Erlangen, with rotation speeds of 0.25 or 0.5 sec. Detector collimation varied from 128x0.6mm single source to 98x0.6x2 and 144x0.4x2 dual source configurations. Spiral pitch factors ranged from 0.3 to 1.3. The mean reference tube current was set at 200 mAs, adjustable to 120 mAs. Automated tube voltage adaptation and tube current modulation were implemented in all instances. Contrast agent administration was standard practice, with an injection rate of 3-4 mL/s and a body weight-adjusted dosage of 400 mg(iodine)/kg (equivalent to 1.14 ml/kg Iomeprol 350mg/ml). All images underwent reconstruction using soft convolution kernels and iterative techniques. Ethical Approval and Data Usage Agreement The data collected for the generation of the datasets involved in this challenge has been approved by an ethical committee (number 23-243-B) held at the Universitätsklinikum Erlangen Hospital. The data to be used during and after the challenge is pseudonymized and coded by the Hospital to assure that a re-identification of the data sample is not possible. Moreover, the patient information is only known by the IP of the Hospital so that the challenge collaborators do not have as well any means to identify patient's data at any point. The data usage agreement for this challenge is CC BY-NC (Attribution-NonCommercial). References [1] Wasserthal, J., Breit, H.-C., Meyer, M. T., Pradella, M., Hinck, D., Sauter, A. W., Heye, T., Boll, D. T., Cyriac, J., Yang, S., Bach, M., & Segeroth, M. (2023). TotalSegmentator: Robust Segmentation of 104 Anatomic Structures in CT Images. Radiology: Artificial Intelligence, 4(4), 230024. https://doi.org/10.1148/ryai.230024 [2] Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211. [3] Rädsch, T., Reinke, A., Weru, V. et al. Labelling instructions matter in biomedical image analysis. Nat Mach Intell 5, 273-283 (2023). https://doi.org/10.1038/s42256-023-00625-5 [4] Heimann T, van Ginneken B, Styner MA, Arzhaeva Y, Aurich V, Bauer C, Beck A, Becker C, Beichel R, Bekes G, Bello F, Binnig G, Bischof H, Bornik A, Cashman PM, Chi Y, Cordova A, Dawant BM, Fidrich M, Furst JD, Furukawa D, Grenacher L, Hornegger J, Kainmüller D, Kitney RI, Kobatake H, Lamecker H, Lange T, Lee J, Lennon B, Li R, Li S, Meinzer HP, Nemeth G, Raicu DS, Rau AM, van Rikxoort EM, Rousson M, Rusko L, Saddi KA, Schmidt G, Seghers D, Shimizu A, Slagmolen P, Sorantin E, Soza G, Susomboon R, Waite JM, Wimmer A, Wolf I. Comparison and evaluation of methods for liver segmentation from CT datasets. IEEE Trans Med Imaging. 2009 Aug;28(8):1251-65. doi: 10.1109/TMI.2009.2013851. Epub 2009 Feb 10. PMID: 19211338. [5] Brachmann, Franz Xaver. Evaluation einer semi-automatischen Segmentierungsmethode für die Volumetrie des renalen Kortex, der Medulla und des gesamten Nierenparenchyms in nativen T1-gewichteten MR-Bildern. PhD thesis, Medizinischen Fakultät def Friedrich-Alexander-Universität Eralngen-Nürnberg, 2021. https://open.fau.de/items/cb6cd403-3178-4884-98c4-5098178658ba [6] Westenberger, Jasmin Barbara. Automatische Gewebesegmentierung der Nieren und des Pankreas im Ganzkörper-MRT mittels Deep Learning. PhD thesis, Eberhard Karls Universität Tübingen, 2021. http://dx.doi.org/10.15496/publikation-63135 The challenge has been co-funded by Proyectos de Colaboración Público-Privada (CPP2021-008364), funded by MCIN/AEI, and the European Union through the NextGenerationEU/PRTR.

临床问题 在医学影像领域,深度学习(Deep Learning, DL)模型常被用于在复杂解剖结构中勾勒目标结构或异常区域,例如肿瘤、血管或器官。由于这些结构本身的复杂性与变异性,精准界定其边界存在固有不确定性;而不同医学专家对真实边界的认知存在差异,即评分者间变异性(interrater variability),进一步加剧了这种不确定性。深度学习模型需应对这些分歧,导致不同标注者生成的分割结果存在不一致,进而可能对诊断与治疗决策产生影响。 针对医学分割任务中深度学习模型的评分者间变异性问题,亟需开发能够捕捉并量化不确定性的鲁棒算法,同时标准化标注流程、推动医学专家间协作以降低变异性,提升基于深度学习的医学影像分析可靠性。评分者间变异性是医学影像分割深度学习领域面临的重大挑战。 此外,实现模型校准——可靠预测的核心要素之一——在处理多类别与多标注者场景时难度显著提升。模型校准的关键在于确保预测概率与事件真实发生概率相符,从而提升模型可靠性。需注意,即便未明确体现,多类别场景下需考虑类别间交互产生的不确定性。同时,引入多标注者的标注数据进一步增加了复杂性:不同专家的意见差异会带来更广范围的变异性与计算复杂度。 因此,开发能够有效捕捉并量化变异性与不确定性,同时适配多类别、多标注者场景细节的鲁棒算法迫在眉睫。在模型校准、精准分割与处理医学标注变异性之间达成平衡,是基于深度学习的医学影像分析实现成功与可靠的关键。 CURVAS挑战赛目标 基于上述所有原因,我们发起了一项涵盖上述所有考量的挑战赛。本次挑战赛将使用腹部CT扫描数据,每例扫描均包含三位不同专家提供的三份标注,每份标注涵盖三类解剖结构:胰腺、肾脏与肝脏。 本次挑战赛的核心宗旨是基于多标注者信息对模型结果进行评估。评估将分为三个独立环节:首先,开展经典戴斯系数(Dice Score)评估与不确定性研究;其次,进行体积测量评估以提供关键临床信息;最后,开展模型校准性研究。所有评估环节均将基于三份不同的标注数据进行。 如需了解挑战赛更多信息,请访问官网报名参与CURVAS(多标注者多器官分割中的校准与不确定性体积评估,Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation)挑战赛。本次挑战赛将于MICCAI 2024(国际医学图像计算和计算机辅助干预会议)举办。 数据集队列 本次挑战赛的数据集共包含90例CT扫描图像,于2023年8月至2023年10月间在埃尔朗根大学医院前瞻性采集。每份CT扫描涵盖四类标注类别:背景(0)、胰腺(1)、肾脏(2)与肝脏(3)。此外,每例CT扫描均由三位来自不同专家的标注者完成标注,涵盖上述四类结构。 训练阶段队列 将提供属于A组的20例CT扫描及其对应标注。我们鼓励参赛者利用公开的多标注者外部标注数据。为训练集提供少量数据并允许使用公开数据集进行训练,旨在提升挑战赛的包容性,使参赛者可基于任意可获取的数据开发算法。此外,使用本次提供的数据进行训练、使用其他数据进行评估,可提升模型对数据集偏移与其他变异性来源的鲁棒性。 验证阶段队列 将使用属于A组的5例CT扫描作为验证阶段数据集。 测试阶段队列 将使用65例CT扫描开展评估,其中包括属于A组的20例、属于B组的22例与属于C组的23例CT扫描。 验证与测试阶段的CT扫描数据集将在挑战赛结束前不予公开。此外,每例CT扫描所属组别信息也将在挑战赛结束后予以公布。 临床规格 纳入标准:最多存在10个直径小于2.0cm的囊肿。此外,存在明显伪影(如呼吸伪影)或配准不完整的CT扫描将被排除。 参与者需年满18周岁,并需提供口头与书面知情同意,允许其CT扫描图像用于本次挑战赛。研究专属同意与宽泛同意均可被接受。在90例患者中,男性51例,女性39例,年龄范围37~94岁,平均年龄65.7岁。所有患者均在德国巴伐利亚州埃尔朗根大学医院接受治疗。未设置额外筛选标准,以确保样本为典型患者队列的代表性集合。 本数据集共包含90例CT扫描,分为三个组别: · A组:囊肿数量≤2个且无轮廓改变病变——共45例CT扫描 · B组:囊肿数量3~5个且无轮廓改变病变——共22例CT扫描 · C组:囊肿数量6~10个且合并部分病变(肝转移瘤、肾积水、肾上腺转移瘤、肾脏缺失)——共23例CT扫描 无论何种情况,参赛者均无法得知每例扫描所属组别。该信息将与完整数据集一同在挑战赛结束后公布。 标注协议 标注生成的第一步为使用TotalSegmentator工具生成粗略标注[1][2]。随后,标注结果将被发送至三位放射科医师(R1、R2、R3),由其修正自动标注并补充可能遗漏的器官。三位标注放射科医师中的一位为医学博士在读博士生,其预先定义了本次数据集队列的纳入排除标准以及解剖实质与非实质的界定规则,并将该规则提供给另外两位标注医师,以确保标注的一致性[3]。另有两位临床医师(C1、C2)对该医学博士在读博士生定义的队列标准进行监督,但未参与标注工作,因此不会对三位放射科医师的标注结果引入偏倚。 本次挑战赛的每类标注结构均配有专属标注指南,以下按器官分别说明: 肝脏:总体而言,肝脏定义为“包含所有内部结构(如血管系统、肿瘤等)的全部肝组织”[4]。因此,门静脉主干本身无需勾勒,门静脉的两大主要分支亦需排除。而后续所有分支均需纳入标注。“若存在部分被包裹的情况(如大血管如腔静脉与门静脉进入或离开肝脏时),被肝组织包裹的部分需纳入分割范围,以此形成肝脏形态的凸包”[4]。侵入肝脏的脂肪组织需排除,胆囊无需标注。扩张尤其是病理性扩张的胆管需纳入肝脏分割范围。 肾脏:需分割左肾与右肾。分割范围包括肾实质,包含肾髓质。需排除肾盂[5]与输尿管,以避免尿液淤积对原始体积测量造成影响。 胰腺:分割胰腺时,无需区分胰头、胰体与胰尾。此外,脾静脉与肠系膜静脉均无需纳入分割范围[6]。但需完整追踪并标注走行全程的胰腺组织。 技术规格 本次使用的CT扫描均为门静脉期增强CT,扫描层厚为0.6~1mm的薄层图像。胸腹CT扫描均采集于患者住院期间,用于满足各类临床诊疗需求。鉴于本次挑战赛聚焦腹部器官,采用Br40软组织卷积核。CT扫描均在埃尔朗根大学医院使用西门子(SIEMENS)CT扫描仪完成,扫描旋转速度为0.25秒或0.5秒。探测器准直范围涵盖128×0.6mm单源配置至98×0.6×2与144×0.4×2双源配置。螺旋螺距因子范围为0.3~1.3。平均参考管电流设置为200mAs,可下调至120mAs。所有扫描均采用自动管电压适配与管电流调制技术。对比剂注射为标准流程:注射速率3~4mL/s,按照体重调整剂量为400mg(碘)/kg(等效于1.14mL/kg的Iomeprol 350mg/mL)。所有图像均采用软组织卷积核与迭代重建技术进行重建。 伦理批准与数据使用协议 本次挑战赛相关数据集的采集已获得埃尔朗根大学医院伦理委员会批准(批准号:23-243-B)。挑战赛期间及赛后使用的数据均已完成去标识化处理,由医院进行编码,确保无法通过数据样本重新识别患者身份。此外,仅医院的项目负责人可接触患者信息,挑战赛合作方在任何阶段均无法识别患者数据。 本次挑战赛的数据使用协议为CC BY-NC(署名-非商业性使用)。 参考文献 [1] Wasserthal, J., Breit, H.-C., Meyer, M. T., Pradella, M., Hinck, D., Sauter, A. W., Heye, T., Boll, D. T., Cyriac, J., Yang, S., Bach, M., & Segeroth, M. (2023). TotalSegmentator: 104种解剖结构在CT图像中的鲁棒分割. Radiology: Artificial Intelligence, 4(4), 230024. https://doi.org/10.1148/ryai.230024 [2] Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net:一种用于基于深度学习的生物医学图像分割的自配置方法. Nature methods, 18(2), 203-211. [3] Rädsch, T., Reinke, A., Weru, V. et al. 标注规则在生物医学图像分析中的重要性. Nat Mach Intell 5, 273-283 (2023). https://doi.org/10.1038/s42256-023-00625-5 [4] Heimann T, van Ginneken B, Styner MA, Arzhaeva Y, Aurich V, Bauer C, Beck A, Becker C, Beichel R, Bekes G, Bello F, Binnig G, Bischof H, Bornik A, Cashman PM, Chi Y, Cordova A, Dawant BM, Fidrich M, Furst JD, Furukawa D, Grenacher L, Hornegger J, Kainmüller D, Kitney RI, Kobatake H, Lamecker H, Lange T, Lee J, Lennon B, Li R, Li S, Meinzer HP, Nemeth G, Raicu DS, Rau AM, van Rikxoort EM, Rousson M, Rusko L, Saddi KA, Schmidt G, Seghers D, Shimizu A, Slagmolen P, Sorantin E, Soza G, Susomboon R, Waite JM, Wimmer A, Wolf I. 基于CT数据集的肝脏分割方法对比与评估. IEEE Trans Med Imaging. 2009 Aug;28(8):1251-65. doi: 10.1109/TMI.2009.2013851. Epub 2009 Feb 10. PMID: 19211338. [5] Brachmann, Franz Xaver. 肾脏皮质、髓质及全肾实质体积测量的半自动化分割方法评估:基于原生T1加权MR图像. 博士论文, 弗里德里希-亚历山大埃尔朗根-纽伦堡大学医学院, 2021. https://open.fau.de/items/cb6cd403-3178-4884-98c4-5098178658ba [6] Westenberger, Jasmin Barbara. 基于深度学习的全身MRI图像中肾脏与胰腺组织自动分割. 博士论文, 蒂宾根大学, 2021. http://dx.doi.org/10.15496/publikation-63135 资助信息 本次挑战赛由MCIN/AEI资助的公私合作项目(Proyectos de Colaboración Público-Privada, CPP2021-008364)以及欧盟通过NextGenerationEU/PRTR项目联合资助。

提供机构:
Zenodo
创建时间:
2024-05-17
二维码
社区交流群
二维码
科研交流群
商业服务