遇见数据集

Hazelnut X-ray classification benchmark: images, labels and reference results

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

This dataset supports a deep-learning benchmark for automated binary classification of hazelnut kernel quality (healthy vs. defective) from X-ray images. The accompanying journal article is listed under Related identifiers on this record. Contents · 799 single-kernel X-ray images — segmented 224 × 224 pixel grayscale PNGs (lossless), grouped into 101 acquisition units (each corresponding to one X-ray scan of up to eight kernels). · Two annotation CSVs — the initial annotation condition (138 healthy / 661 defective kernels) and the reassessed annotation condition (153 healthy / 646 defective, following expert re-examination of 15 uncertain kernels). Both CSVs reference the same image files. · Reproducibility code (MIT) — the group-wise split-rotation benchmark script that trains seven single-model configurations (a custom CNN with three loss variants, ImageNet-pretrained Swin Transformer Tiny in fully-fine-tuned and frozen variants, EfficientNet-B0 and ResNet-18) and ten probability-aggregation ensembles across five stratified data splits, plus a supplementary nut-level split variant. · Reference results (JSON) — the exact numerical output of the benchmark on the reference hardware (NVIDIA A100-40GB, PyTorch 2.10, CUDA 12.8), sufficient for bit-exact verification of reproduction. Key result Under the reassessed annotation condition, the average-probability ensemble combining a BCE-trained convolutional neural network with a frozen Swin Transformer Tiny achieved a mean balanced accuracy of 86.26 ± 1.80 % across five stratified data splits. Data provenance Kernels of Corylus avellana var. pontica cv. Anakliuri were harvested in Zugdidi (Georgia, 2024) and radiographed with a MILabs U-CT scanner (65 kV, 0.25 mA, 500 µm Al filtration). Kernel labels come from a two-stage physical inspection: external UNECE grading, followed by destructive sectioning of externally-healthy kernels to identify internal defects. Reproduction The main benchmark runs end-to-end in ~50 minutes on a single A100 GPU. See README.md inside the archive for full file structure, schemas, verification procedure, and design choices. Licensing Data and reference results are released under CC-BY 4.0; code under MIT (see LICENSE-data.txt and LICENSE-code.txt).

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务