遇见数据集

SapBark-64: A dataset of bark images for 64 fruit-tree sapling classes

收藏
Zenodo2026-02-27 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is described in the following article: SapBark-64: A dataset of bark images for 64 fruit-tree sapling classes. Data in Brief, 64, 112354. https://doi.org/10.1016/j.dib.2025.112354 SapBark-64, a dataset of sapling bark images covering 64 fruit tree classes (species/cultivars), contains a total of 5,742 close-up photographs. The images were taken in 2025 from 1–2-year-old saplings available for sale at three commercial nurseries in Trabzon. A separate photo of the nursery tag was also taken for each class to support verification. The images were taken with an iPhone 16 Pro Max, approximately 10 cm from the trunk, under a white background to mask scene clutter, and under as uniform natural lighting as possible. This layout ensures the preservation of fine-scale bark morphology and ensures comparability between images. The dataset contains 57–149 images per class. The repository structure is kept simple to facilitate direct research: (i) 64 class folders for raw images (JPG) and (ii) background-removed images (WebP), each named by cultivar name and matching one-to-one; (iii) a structured Excel (XLSX) metadata file containing standard class-level fields (family, scientific/common name, cultivar/variety, sapling height, trunk diameter, recommended planting season, growth rate, fruiting age, average yield per tree, production region, propagation method, etc.). The background-removed directory mirrors the raw image hierarchy; matching can be done directly using class folders and the same base file name if necessary. The dataset is designed for both computer vision tasks such as fine-grained classification and content retrieval, as well as feature-based analyses relating bark texture to horticultural traits. A consistent labeling scheme, recommended training/validation/test distinctions, and file naming conventions have been adopted for this purpose. The controlled acquisition protocol (≈10 cm distance, white background, uniform illumination) is clearly documented. Subtle cultivar differences and a detailed label space of 64 classes also provide a powerful testbed for studies on few-sample learning, domain adaptation, open-set recognition, and robustness. For background removal, a fully automated Python process based on Fast Segment Anything (FastSAM) was used. Point-guided masking focused on the trunk; assuming the trunk occupies a smaller area than the background, the smallest-area mask was selected. Automatic inversion was performed if the model inadvertently highlighted the background. The process resulted in images with transparent backgrounds; the file structure and class labels were preserved. Label photos for each class were removed from the new index by visual inspection. The dataset covers only cultivated saplings; no human or animal data is included. The images were taken with the permission of the nursery owners and under the supervision of an agronomist; no personally identifiable information is included; and the geographic detail is limited to the city level (Trabzon, Türkiye). Therefore, institutional ethics approval/informed consent was not required. SapBark-64 is available for publication on Zenodo under an open license (e.g., CC BY 4.0) with a DOI. The acquisition conditions, indexing, and metadata schema are documented to ensure clear and reliable reuse. The dataset aims to directly contribute to the computer vision, plant phenotyping, and horticulture communities, as well as nursery practitioners (label verification, traceability, decision support). In conclusion, this dataset, focusing on bark texture, a relatively stable biomarker that is observable throughout the year, fills a critical gap in early-stage sapling identification and provides a reproducible, comparable, and extensible foundation for both academic and practical workflows.

提供机构:
Zenodo
创建时间:
2025-10-06
二维码
社区交流群
二维码
科研交流群
商业服务