BoneMarrowWSI-PediatricLeukemia: A Comprehensive Dataset of Bone Marrow Aspirate Smear Whole Slide Images with Expert Annotations and Clinical Data in Pediatric Leukemia
收藏资源简介:
This dataset corresponds to a collection of images and/or image-derived data available from the National Cancer Institute Imaging Data Commons (IDC). This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using the IDC Portal. You can use the manifests included in this Zenodo record to download the collection following the Download instructions below. Collection Description Image data: The dataset comprises bone marrow aspirate smear WSI for 245 pediatric cases (< 18 years) of leukemia, including acute lymphoid leukemia (ALL), acute myeloid leukemia (AML), and chronic myeloid leukemia (CML). The smears were prepared for the initial diagnosis (i.e., without prior treatment), stained in accordance with the Pappenheim method, and scanned at 40x magnification (without immersion), resulting in a resolution of 0.11x0.11 µm/pixel. Metadata: Additionally, clinical information (age group, sex, diagnosis) and laboratory data (blasts, white blood cell count, thrombocytes, LDH, uric acid, hemoglobin) are available for each case. Annotations: The images have been annotated with rectangular regions of interest (ROI) within the evaluable monolayer area, and a total of 47176 cell bounding box annotations have been placed within the regions of interest. Cells have been annotated by multiple experts in a consensus labeling approach with 49 distinct cell type classes. This consensus approach entailed that each cell was sequentially annotated by multiple individuals until each cell had been labeled by at least two individuals, and the majority class was assigned in at least half of all annotations for that image. The labels from all annotation sessions, as well as the final consensus class for each cell, are available as DICOM Bulk Annotations (ANN modality) series in the collection. Images were originally obtained as proprietary MIRAX files using 3DHistech scanners, but were afterwards converted to standard DICOM format by the IDC team. Clinical data are contained in the DICOM metadata. In addition lab values are available as BigQuery table. The accompanying preprint describes the study and the dataset in detail. Conversion into DICOM was done using the scripts in https://github.com/ImagingDataCommons/conversion_mirax_dicom, which rely on the wsidicomizer library. A detailed description of the conversion of images and annotations into DICOM is appended to this Zenodo record as a pdf (Documentation_of_conversion_process.pdf). (!) Please note that due to technical issues one slide of the reported 246 slides (Slide ID: C45147DEE729311EF5B5C3003946C48F_1_bm and it's accompanying annotations ( 2 ROIs and 64 cell annotations with labels) could not be included into IDC. (!) References Höfener, H., Kock, F., Pontones, M., Ghete, T., Pfrang, D., Dickel, N., Kunz, M., Schacherer, D. P., Clunie, D. A., Fedorov, A., Westphal, M. & Metzler, M. From data to diagnosis: A large, comprehensive bone marrow dataset and AI methods for childhood leukemia prediction. arXiv [cs.LG] (2025). at <http://arxiv.org/abs/2509.15895> Acknowledgments The authors thank Stefanie Barnickel, Nathalie Dollmann, Tatjana Flamann, Meinolf Suttorp, and Perdita Weller for the labelling of the cells. The authors thank the following institutions for supplying BMA smears: University Hospital Augsburg (Univ.-Prof. Dr. Dr. med. Michael Frühwald), Charité Berlin - ALL-REZ BFM Study Group (PD Dr. med. Arend von Stackelberg), University Hospital at the TU Dresden (Prof. Dr. med. Meinolf Suttorp), University Hospital Essen - AML-BFM Study Group (Prof. Dr. Dirk Reinhardt), Technical University of Munich (Prof. Dr. med. Irene Teichert-von Lüttichau), University Hospital Würzburg (Prof. Dr. med. Matthias Eyrich). This study was supported by a grant from the German Federal Ministry of Education and Research (FKZ: 031L0262A; BMDeep) Preparation of the Dataset for publication was partly supported by Federal funds from the National Cancer Institute, National Institutes of Health (Task Order No. HHSN26110071 under Contract HHSN261201500003l). The entire dataset is made available in National Cancer Institute Imaging Data Commons (https://imaging.datacommons.cancer.gov). If you have any questions about the dataset please contact IDC support at support@canceridc.dev. Files included A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example, bonemarrowwsi_pediatricleukemia-idc_v22-aws.s5cmd corresponds to the contents of the bonemarrowwsi_pediatricleukemia collection introduced in IDC data release v22. bonemarrowwsi_pediatricleukemia-idc_v24-aws.s5cmd: AWS download manifest bonemarrowwsi_pediatricleukemia-idc_v24-gcs.s5cmd: GCS download manifest bonemarrowwsi_pediatricleukemia-idc_v24-dcf.dcf: DCF download manifest Manifest files ending in -aws.s5cmd reference files in Amazon Web Services (AWS) buckets; -gcs.s5cmd reference files in Google Cloud Storage. The actual files are identical and mirrored between AWS and GCP. Download instructions Each manifest file includes instructions in its header on how to download the included files. To download the files using .s5cmd manifests: Install idc-index: pip install --upgrade idc-index Download the files referenced by a manifest included in this dataset: idc download manifest.s5cmd To download files using a .dcf manifest, see the manifest header. For questions or help, contact support@canceridc.dev or post on the IDC Forum.



