Mitosis Subtyping Dataset
收藏资源简介:
Overview This dataset is a curated subset of the final dataset developed for atypical mitosis detection in histopathology whole slide images (WSIs). It contains 8,236 image patches of size 64×64 pixels, categorized into four classes: Artifact Mimicker Mitosis Atypical Mitosis The dataset is split into three cross-validation folds, allowing for standardized evaluation across multiple training/testing splits. Dataset Generation Process To construct a high-quality dataset for training and evaluating atypical mitosis detection algorithms, we employed an AI-assisted labeling pipeline: Initial Model TrainingAn initial Atypical Mitosis Detection (AMD) model was trained using a combination of publicly available datasets — MIDOG21 and TUPAC — as described by Fick, Bertram, and Aubreville (2024), along with an internal dataset containing artifacts. These training sources provided approximately 9,500 patches of 64×64 pixels at 40x magnification, covering four target categories: artifact, mimicker, mitosis, and atypical mitosis. To access this dataset, please directly contact Marc Aubreville. Diverse Patch Sampling from TCGATo enhance diversity and representativeness, 6,000 image patches were sampled from The Cancer Genome Atlas (TCGA). For 40x WSIs, 64×64 patches were directly extracted. For 20x WSIs, 32×32 patches were extracted and then resized to 64×64 pixels.Feature embeddings of these patches were generated using the UNI foundation model. Clustering for DiversityThe extracted embeddings were clustered into six distinct groups using the K-Nearest Neighbors algorithm, with the number of clusters chosen using the Elbow Method. From each cluster, 1,000 patches were randomly selected, totaling 6,000 diverse candidates. AI-Assisted LabelingThese 6,000 patches were pre-classified using the initially trained AMD model. The predictions were then manually reviewed and corrected by three board-certified pathologists to ensure high labeling accuracy. Final Dataset CompositionThe validated 6,000 TCGA-derived patches were merged with some previously labelled patches across the four diagnostic categories, resulting in this release (8,236 samples). Use Cases and Applications This dataset is intended for use in developing and benchmarking machine learning models for mitosis detection and classification in digital pathology. Potential applications include: Training deep learning models for mitotic figure classification. Evaluating performance on atypical mitosis detection.



