遇见数据集

Pixel-level Protected Health Information (PHI) - Supplement to Exploring AI-Based System Design for Pixel-Level Protected Health Information Detection in Medical Images

收藏
Zenodo2025-09-25 更新2026-05-26 收录
官方服务:

资源简介:

If you plan to use this dataset, please cite the following paper: Truong, T., Baltruschat, I.M., Klemens, M. et al. Exploring AI-Based System Design for Pixel-Level Protected Health Information Detection in Medical Images. J Digit Imaging. Inform. med. (2025). https://doi.org/10.1007/s10278-025-01619-y This dataset includes two collections: RadPHI-test and MIDI.RadPHI-test includes 1000 images across four modalities: CT, chest X-ray, radionuclide bone scan, and MRI images overlaid with synthetic texts. Images are sourced from the following datasets: TotalSegmentator [1] for CT, BS-80K [2] for bone scans, ChestX-ray8 [3] for chest X-rays, and BRATS[4] for brain MRI. The imprints are synthetically generated over 16 categories, six of which are considered PHI: patient name, address, identifier, phone number, email, and date. Of the 1000 images, 850 contain at least one type of PHI imprint.MIDI is curated from the validation and test set of the 2024 Medical Image De-Identification Benchmark (MIDI-B) challenge [5], which is available on The Cancer Imaging Archive [6]. This dataset originally consists of 605 studies across multiple modalities, each containing synthetic PHI content embedded at both the DICOM header and pixel level. We utilize a DICOM viewer, specifically MD.ai [7], to overlay the DICOM tags onto the images. We randomly sample DICOM tags to ensure that the generated imprints represent all possible PHI categories, similar to the RadPHI-test dataset. After applying the overlays, we export the images from the viewer. The resulting images may include not only the DICOM tag overlays but also burn-ins by the challenge organizers. The final version of the dataset comprises 550 images categorized into five PHI types: patient name, address, identifier, phone number, and date. We performed instance-level annotation of the images by generating coordinates for PHI instances along with their corresponding categories. This annotation process was carried out and validated by two independent annotators to ensure accuracy and reliability. [1] Wasserthal J, Breit HC, Meyer MT, Pradella M, Hinck D, Sauter AW, Heye T, Boll DT, Cyriac J, Yang S, et al.: TotalSegmentator: robust segmentation of 104 anatomic structures in CT images. Radiol Artif Intell 5(5), 2023. [2]Huang Z, Pu X, Tang G, Ping M, Jiang G, Wang M, Wei X, Ren Y: BS-80K: The first large open-access dataset of bone scan images. Comput Biol Med 151:106221, 2022. [3] Wang X, Peng Y, Lu L, Lu Z, Bagheri M, Summers RM: ChestX-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2097–2106, 2017. [4]Antonelli M, Reinke A, Bakas S, Farahani K, Kopp-Schneider A, Landman BA, Litjens G, Menze B, Ronneberger O, Summers RM, et al.: The medical segmentation decathlon. Nat Commun 13(1):4128, 2022. [5] Farahani K, Clunie D, Klenk J, Kopchick B, Diaz M, Pan Q, Pei L, Prior F, Rutherford M, Singh A, Sutton G, Wagner U: Medical Image De-Identification Benchmark (MIDI-B). Available at https://www.synapse.org/Synapse:syn53065760 Accessed 16 April 2025. [6] Rutherford MW, Nolan T, Pei L, Wagner U, Pan Q, Farmer P, Smith K, Kopchick B, Opsahl-Ong L, Sutton G, Clunie DA, Farahani K, Prior F: Data in support of the MIDI-B Challenge (MIDI-B-Synthetic-Validation, MIDI-B-Curated-Validation, MIDI-B-Synthetic-Test, MIDI-B-Curated-Test) (Version 1) [Data set]. The Cancer Imaging Archive, https://doi.org/10.7937/cf2p-aw56, 2025 [7] MD.ai. Available at https://www.md.ai. Accessed 28 April 2025.

提供机构:
Zenodo
创建时间:
2025-07-29
二维码
社区交流群
二维码
科研交流群
商业服务