2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 3,001-4,000
收藏资源简介:
This upload contains slices 3,001 – 4,000 from the data collection described in Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka “"2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning", <em>Sci Data</em> <strong>10</strong>, 576 (2023) or arXiv:2306.05907 (2023) Abstract:<br> "Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline." The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by TESCAN-XRE NV, located in the FleX-ray Lab at the Centrum Wiskunde & Informatica (CWI) in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels, \(74.8\mu m^2\) each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another. Please refer to the paper for all further technical details. The complete dataset can be found via the following links: 1-1000, 1001-2000, 2001-3000, 3001-4000, 4001-5000, OOD.<br> The reference reconstructions and segmentations can be found via the following links: 1-1000, 1001-2000, 2001-3000, 3001-4000, 4001-5000, OOD. The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on github. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well. Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc “export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE”. For more information or guidance in using the data collection, please get in touch with Maximilian.Kiss [at] cwi.nl Felix.Lucka [at] cwi.nl
本上传文件包含来自Maximilian B. Kiss、Sophia B. Coban、K. Joost Batenburg、Tristan van Leeuwen与Felix Lucka发表于*《Sci Data》*2023年第**10**卷第576页的论文"2DeteCT——面向机器学习的大型可扩展二维可训练实验计算机断层扫描数据集",以及arXiv:2306.05907 (2023)预印本中所述数据集的3001至4000号切片。 **摘要**:计算成像领域的近期研究多聚焦于开发用于图像重建的机器学习(ML, Machine Learning)技术,而此类技术的训练需要包含测量数据与真值图像的大规模数据集。然而,适用于X射线计算机断层扫描(CT, Computed Tomography)的实验数据集十分稀缺,现有方法往往仅在模拟数据上开发与评估。为填补这一空白,我们向社区开放了一款通用型二维扇形束CT数据集,可用于开发面向多种图像重建任务的ML技术。本数据集通过一套高度灵活的实验室X射线CT装置,采用经精心设计的半自动扫描流程采集得到。我们对一批形状与密度自然差异显著的多样化样本逐切片进行扫描(总计5000个切片),扫描具备高角度与空间分辨率,并采用三种不同的射线束特性模式:高保真模式、低剂量模式与诱发束硬化模式。此外,我们还采集了750个分布外(OOD, Out-of-Distribution)切片,通过调整样本与射线束参数,以适配鲁棒性与图像分割任务。本数据集包含原始投影数据、参考重建图像以及基于开源数据处理流程生成的分割掩码。 本数据集采集自荷兰阿姆斯特丹数学与计算机科学中心(CWI, Centrum Wiskunde & Informatica)FleX-ray实验室中由TESCAN-XRE NV研发的FleX-ray型定制CT扫描仪,该设备具备高度灵活性与可编程性。其核心组件包括一台锥束微焦点X射线点光源(额定参数:最高90kV、90W),可发出多色X射线并投射至搭载碘化铯(Tl)闪烁体的14位CMOS(互补金属氧化物半导体,complementary metal-oxide semiconductor)平板探测器(型号Dexella 1512NDT,分辨率为1536×1944像素,单像素面积74.8μm²)。为构建二维数据集,我们仅读取探测器的中央行像素,以模拟扇形束几何布局。在射线源与探测器之间设有旋转载物台,用于放置扫描样本。设备的三大核心组件(射线源、探测器面板与旋转载物台)均安装于独立平移滑轨,可实现各组件的独立运动。更多技术细节请参阅原论文。 完整数据集可通过以下链接获取:1-1000、1001-2000、2001-3000、3001-4000、4001-5000以及分布外切片集。参考重建图像与分割掩码可通过以下链接获取:1-1000、1001-2000、2001-3000、3001-4000、4001-5000以及分布外切片集。与论文中所述投影数据加载、预处理、重建与分割流程对应的Python脚本可在GitHub仓库中获取。同时,该仓库还包含一份机器可读文件,记录了各采集模式下的扫描参数与仪器数据,以及对应的加载脚本。 **注意事项**:解压.zip压缩包时建议使用图形用户界面。若在Linux系统中解压时触发zip炸弹检测错误,可通过设置`UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE`环境变量解决:可在.bashrc文件中添加`export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE`后重新运行解压命令。如需获取数据集使用的更多信息或指导,请联系Maximilian.Kiss [at] cwi.nl或Felix.Lucka [at] cwi.nl。



