ChestX6: A Compiled Six-Class Chest X-Ray Classification Benchmark
收藏资源简介:
ChestX6 is a curated, six-class chest X-ray image dataset (18,036 images) compiled from four established, publicly available chest X-ray sources. It is a compilation, not a single-site acquisition — per-class provenance is documented below for full transparency. It is released as a general-purpose resource for chest X-ray classification and self-supervised pretraining research. CLASS COMPOSITION AND PROVENANCE Class Count Source License/Attribution COVID-19 3,017 COVID-19 Radiography Database (Chowdhury et al.) Free to use; attribution required Normal 3,271 COVID-19 Radiography Database (Chowdhury et al.) Free to use; attribution required Pneumonia-Bacterial 3,000 Chest X-Ray Images (Pneumonia), Kermany et al. CC BY 4.0 Pneumonia-Viral 3,013 Chest X-Ray Images (Pneumonia), Kermany et al. CC BY 4.0 Emphysema 2,550 NIH ChestX-ray14 (Emphysema-label subset), Wang et al. Public domain; attribution requested Tuberculosis 3,185 Tuberculosis (TB) Chest X-ray Database, Rahman et al. Derived exclusively from the freely available subset; attribution required Total: 18,036 images. This release has not been deduplicated; a small number of near-duplicate and exact-duplicate images are known to exist within it (see Tuberculosis note below) — apply whatever deduplication or splitting approach suits your own use case. TUBERCULOSIS CLASS NOTE: the source TB database contains a freely available subset of approximately 700–703 original TB-positive chest X-ray images. The 3,185 tuberculosis images included in ChestX6 were generated exclusively from this freely available subset through data augmentation. Thus, the 3,185 images do not represent 3,185 independent original X-ray acquisitions. The source collection also contains approximately 2,800 additional original TB images that are available through the NIAID TB Portal under a separate data-use agreement. None of these additional restricted images were used or included in ChestX6. Because the 3,185 TB images are derived from a relatively small pool of original images, multiple images may share the same source image. Researchers creating their own train/validation/test splits should account for these related images and, where possible, keep images derived from the same original source image within the same split to avoid data leakage. FILES IN THIS RELEASE: images/ (18,036 files, organized by class), README.md. CITING: please cite the four original source datasets listed in the provenance table above, per their respective attribution requirements.



