SapBark-64: A dataset of bark images for 64 fruit-tree sapling classes
收藏资源简介:
This dataset is described in the following article: SapBark-64: A dataset of bark images for 64 fruit-tree sapling classes. Data in Brief, 64, 112354. https://doi.org/10.1016/j.dib.2025.112354 SapBark-64, a dataset of sapling bark images covering 64 fruit tree classes (species/cultivars), contains a total of 5,742 close-up photographs. The images were taken in 2025 from 1–2-year-old saplings available for sale at three commercial nurseries in Trabzon. A separate photo of the nursery tag was also taken for each class to support verification. The images were taken with an iPhone 16 Pro Max, approximately 10 cm from the trunk, under a white background to mask scene clutter, and under as uniform natural lighting as possible. This layout ensures the preservation of fine-scale bark morphology and ensures comparability between images. The dataset contains 57–149 images per class. The repository structure is kept simple to facilitate direct research: (i) 64 class folders for raw images (JPG) and (ii) background-removed images (WebP), each named by cultivar name and matching one-to-one; (iii) a structured Excel (XLSX) metadata file containing standard class-level fields (family, scientific/common name, cultivar/variety, sapling height, trunk diameter, recommended planting season, growth rate, fruiting age, average yield per tree, production region, propagation method, etc.). The background-removed directory mirrors the raw image hierarchy; matching can be done directly using class folders and the same base file name if necessary. The dataset is designed for both computer vision tasks such as fine-grained classification and content retrieval, as well as feature-based analyses relating bark texture to horticultural traits. A consistent labeling scheme, recommended training/validation/test distinctions, and file naming conventions have been adopted for this purpose. The controlled acquisition protocol (≈10 cm distance, white background, uniform illumination) is clearly documented. Subtle cultivar differences and a detailed label space of 64 classes also provide a powerful testbed for studies on few-sample learning, domain adaptation, open-set recognition, and robustness. For background removal, a fully automated Python process based on Fast Segment Anything (FastSAM) was used. Point-guided masking focused on the trunk; assuming the trunk occupies a smaller area than the background, the smallest-area mask was selected. Automatic inversion was performed if the model inadvertently highlighted the background. The process resulted in images with transparent backgrounds; the file structure and class labels were preserved. Label photos for each class were removed from the new index by visual inspection. The dataset covers only cultivated saplings; no human or animal data is included. The images were taken with the permission of the nursery owners and under the supervision of an agronomist; no personally identifiable information is included; and the geographic detail is limited to the city level (Trabzon, Türkiye). Therefore, institutional ethics approval/informed consent was not required. SapBark-64 is available for publication on Zenodo under an open license (e.g., CC BY 4.0) with a DOI. The acquisition conditions, indexing, and metadata schema are documented to ensure clear and reliable reuse. The dataset aims to directly contribute to the computer vision, plant phenotyping, and horticulture communities, as well as nursery practitioners (label verification, traceability, decision support). In conclusion, this dataset, focusing on bark texture, a relatively stable biomarker that is observable throughout the year, fills a critical gap in early-stage sapling identification and provides a reproducible, comparable, and extensible foundation for both academic and practical workflows.
本数据集的相关描述见于以下文献: 《SapBark-64:面向64类果树幼树的树皮图像数据集》,发表于《Data in Brief》,第64卷,文章编号112354。https://doi.org/10.1016/j.dib.2025.112354 SapBark-64是覆盖64个果树类群(物种/栽培品种)的幼树树皮图像数据集,共包含5742张特写照片。所有图像均拍摄于2025年,采集自土耳其特拉布宗(Trabzon)三家商业苗圃在售的1~2年生幼树。为辅助验证,每一类群还拍摄了对应的苗圃标签照片。图像均使用iPhone 16 Pro Max拍摄,拍摄距离距树干约10厘米,采用白色背景以屏蔽场景杂乱信息,并尽可能使用均匀的自然光照。该拍摄方案可保留树皮的精细形态特征,同时保证不同图像间的可比性。数据集每类群包含57~149张图像。 为便于直接开展研究,数据集的存储结构设计简洁:(i) 64个类群文件夹,分别存放原始图像(JPG格式)与去背景图像(WebP格式),两类图像均以栽培品种名称命名,一一对应;(ii) 结构化的Excel(XLSX格式)元数据文件,包含标准的类群级字段,如科属、学名/通用名、栽培品种、幼树高度、树干直径、推荐种植季节、生长速率、挂果年限、单株平均产量、产区、繁殖方式等。去背景图像目录与原始图像目录层级一致,必要时可直接通过类群文件夹与相同的基础文件名完成图像匹配。 本数据集可用于细粒度分类、内容检索等计算机视觉任务,以及基于树皮纹理与园艺性状关联的特征分析研究。为此,数据集采用了统一的标注方案、推荐的训练/验证/测试集划分规则与文件命名规范。严格受控的采集方案(约10厘米拍摄距离、白色背景、均匀光照)已完整归档。类群间细微的栽培品种差异与64个类别的详细标注空间,也为少样本学习、领域自适应、开放集识别与鲁棒性等研究提供了优质测试平台。 背景移除环节采用了基于Fast Segment Anything(FastSAM)的全自动化Python流程。流程采用点引导掩码生成,聚焦于树干区域;假设树干面积小于背景面积,因此选择面积最小的掩码。若模型误将背景识别为前景,则自动执行掩码反转。该流程最终生成背景透明的图像,且保留原有的文件结构与类群标签。通过人工目视检查,移除了每个类群对应的标签照片索引。 本数据集仅包含栽培幼树的相关数据,未涉及任何人类或动物数据。所有图像均在获得苗圃所有者许可,并经农艺师监督的情况下拍摄,未包含任何个人可识别信息,且地理信息仅精确到城市级别(土耳其特拉布宗市)。因此,本数据集无需机构伦理审查或知情同意。 SapBark-64已通过Zenodo平台以开放许可(如CC BY 4.0)发布,并配有DOI编号。数据集的采集条件、索引规则与元数据架构均已完整归档,以确保可清晰、可靠地复用。本数据集旨在为计算机视觉、植物表型组学与园艺学领域的研究者,以及苗圃从业者(用于标签验证、溯源与决策支持)提供直接支持。 综上,本数据集聚焦于树皮纹理——一种全年可观测的相对稳定的生物标志物,填补了幼树早期识别领域的关键空白,并为学术研究与实际应用流程提供了可复现、可对比、可扩展的基础支撑。



