SlideGuard: WSI Artifact Detection Benchmark Dataset
收藏资源简介:
SlideGuard: WSI Artifact Detection Benchmark Dataset Overview This repository contains 43 raw Whole Slide Image (WSI) files in SVS format required to participate in the SlideGuard Artifact Detection Benchmark. The set includes five slides per type (total of 25) from five tissue types from The Cancer Genome Atlas (breast, colon, lung, prostate and lymphatic tissue), six lymphatic tissue slides from the Maria Sklodowska-Curie National Research Institute of Oncology and twelve mouse skin tissue slides. Please note ten remaining breast slides from Dartmouth-Hitchcock Medical Center must be downloaded independently from the original distribution site as described below. This repository solely provides the raw input data for artifact detection methods. This dataset was curated and released as part of our work currently under review for the MICCAI 2026 Main Conference. --- Benchmark Workflow & Instructions To participate in the benchmark and obtain valid, comparable mean results, you must evaluate your method on all 53 original slides included in our study. 1. Download the raw data. Acquire all 53 WSI files (see *Dataset Composition* below).2. Run your method. Process the WSIs using your artifact detection algorithm.3. Generate compatible predictions. Output your predictions as 8-bit grayscale PNG masks (0=Background, 1=Tissue, 2=Artifact).4. Evaluate on Codabench. Zip your prediction masks and upload them to our live benchmarking platform to receive your metrics: https://www.codabench.org/competitions/16457/ --- Dataset Composition To comply with varying data redistribution policies across our source institutions (detailed in the manuscript), the 53 WSIs must be compiled from two locations: 1. Directly Hosted Data (43 Slides) This Zenodo repository directly hosts **43 WSIs** compiled from three open-access cohorts. You can download these `.svs` files directly from the files section below. 2. The DHMC Cohort (10 Slides) Due to institutional redistribution restrictions, we cannot directly host the WSIs originating from the Dartmouth-Hitchcock Medical Center. To complete the benchmark dataset, you must download these manually:1. Navigate to the official Dartmouth Lung Cancer Histology Dataset download site.2. Proceed to the data download section.3. Select the specific files corresponding to the following IDs used in our benchmark: * `DHMC_0016` * `DHMC_0017` * `DHMC_0042` * `DHMC_0044` * `DHMC_0049` * `DHMC_0055` * `DHMC_0063` * `DHMC_0071` * `DHMC_0086` Once you have combined these DHMC slides with the 43 slides downloaded from this repository, you will have the complete 53-slide dataset required to run the Codabench evaluation. --- Citation If you use this dataset or the Codabench evaluation tool in your research, please cite our manuscript: @unpublished{kaczmarek2026slideguard_submitted, author = {Kaczmarek, Gabriela and Krawczyk-Borysiak, Zuzanna and Miller, Mateusz and Krawczyk, Adam and Sokol, Malgorzata and Przybylska, Martyna and Szymanski, Lukasz and Markiewicz, Tomasz and Swiderska-Chadaj, Zaneta}, title = {SlideGuard: {WSI} Artifact Detection Benchmark}, note = {Manuscript submitted to MICCAI 2026}, year = {2026}}



