遇见数据集

<b>PEMDefect-1107</b>

收藏
Figshare2025-12-28 更新2026-04-08 收录
官方服务:

资源简介:

<b>PEMDefect-1107</b> is a curated grayscale image dataset for <b>single-class defect localization</b> in proton-exchange membrane (PEM) fuel cell materials. The dataset is designed for <b>bounding-box–based defect detection</b> and supports research in materials-focused computer vision, data-centric machine learning, and defect localization under data-scarce conditions.The dataset contains <b>1,107 images in total</b> and is organized under a single root directory with <b>explicit train, validation, and test splits</b>. The splits consist of <b>1,045 training images</b>, <b>31 validation images</b>, and <b>31 test images</b>. Each split is stored in a dedicated subdirectory and follows a <b>split-aware directory structure</b> to prevent data leakage.All bounding-box annotations are provided in <b>YOLO format</b>, with one label file per image. This format ensures a strict one-to-one correspondence between images and annotations and avoids centralized annotation mismatch issues. A dataset configuration file (<code>dataset.yaml</code>) is included to define split paths and class metadata, enabling immediate use in standard object detection pipelines.All images are provided in <b>grayscale</b>, resized to a fixed resolution of <b>640 × 640 pixels</b>, and stored in <b>PNG format</b>. A consistent, physics-aware preprocessing pipeline was applied across the dataset, including conversion from microscopy acquisition formats, deterministic removal of scale bars via cropping, and global intensity normalization. Scale bars were removed to prevent shortcut learning from acquisition metadata and to ensure models learn from intrinsic material structure rather than imaging artifacts.The train/validation/test split was finalized <b>prior to any data augmentation</b>, ensuring strict separation between splits. Data augmentation was applied <b>exclusively to the training set</b>, while validation and test images remain unaugmented, enabling unbiased and reproducible evaluation of defect localization performance.The dataset underwent comprehensive integrity and validation checks, including verification of image–annotation correspondence, geometric consistency of bounding boxes, absence of truncation artifacts, detection of duplicate images, and confirmation of split integrity with no leakage across subsets. All hard validation criteria were satisfied in the final release.PEMDefect-1107 is suitable for training and benchmarking defect localization models such as Faster R-CNN and RetinaNet (via standard annotation conversion), as well as for studying data-centric effects, augmentation strategies, and robustness in materials-oriented computer vision tasks.

创建时间:
2025-12-28
二维码
社区交流群
二维码
科研交流群
商业服务