遇见数据集

A Systematically Designed Open-Access SEM Dataset and Robustness Benchmark for Machine Learning-Based Microstructure Segmentation: A High-Chromium Cast Iron Case Study

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

This record provides an open-access scanning electron microscopy (SEM) dataset of high-chromium cast iron (HCCI, 26 wt.% Cr) microstructures, designed to isolate acquisition-induced image variance from material-induced variance. The dataset comprises 777 micrographs systematically acquired across eight controlled variation axes: four SEM instruments (Helios G4 PFIB CXe, Helios NanoLab, VEGA3 XMH, Zeiss Gemini; FEG and W sources), three detector types (SE, BSE, InLens), three accelerating voltages (5, 10, 20 kV), two etching reagents (Vilella's reagent, Nital), three sample states (as-cast; 980 °C + water quenched; 980 °C for 9 h + air cooled), two image scales (overview and detail), and two dwell-time categories. The defining feature of the dataset is its correlative structure. For each of 12 unique regions of interest (3 specimens × 2 image scales × 2 etching conditions), the identical microstructural area was imaged under as many acquisition conditions as physically permitted, and all micrographs of a stack were registered onto a common reference frame using SIFT-based and manual feature registration in Fiji. Consequently, the underlying microstructure remains fixed while the imaging modality varies systematically, and 12 manually annotated ground-truth masks propagate across the full set of 777 images. Annotation was performed class-wise using trainable WEKA segmentation with subsequent expert correction, exploiting complementary SE and BSE contrast; the label scheme distinguishes up to seven phase classes (austenite, fresh martensite, retained austenite, eutectic carbides, tempered martensite, secondary carbides, and a combined tempered martensite + secondary carbide class where resolution is insufficient), supporting binary, four-class, and seven-class segmentation tasks. Every micrograph is accompanied by a machine-readable metadata record (CSV) containing unique ID and filename, microscope, detector, accelerating voltage, magnification, etching agent, sample, imaging kind, pixel size, beam current, dwell time (numeric and categorical), chamber pressure, and working distance, with all numeric values in SI units. The dataset follows FAIR principles. The correlative design makes the data suitable for uses beyond standard supervised segmentation, including domain-adaptation and domain-generalization benchmarking, multi-detector fusion, metadata-conditioned segmentation, genuinely paired image-to-image translation between detectors and instruments, and paired denoising and super-resolution studies based on the low-/high-dwell-time and overview/detail acquisitions. A segmentation showcase using this dataset, together with a discussion of acquisition-related robustness and suggested directions for further work, is presented in the accompanying publication of the same title [Journal / DOI — to be added].

提供机构:
Zenodo
创建时间:
2026-08-14
二维码
社区交流群
二维码
科研交流群
商业服务