A real-world leaf image dataset of seven Bangladeshi medicinal plant species
收藏资源简介:
This dataset contains a curated collection of leaf images from seven commonly used medicinal plant species native to Bangladesh, developed to support research in plant identification and classification using computer vision and machine learning techniques. The central research hypothesis is that leaf morphology provides sufficient visual information to discriminate medicinal plant species under real-world outdoor conditions. The dataset comprises a total of 13,095 leaf images, including 6,700 raw images and 6,395 pre-processed images. All images are resized to 640 × 640 pixels. However, the images represent seven medicinal plant species widely used in traditional and modern healthcare practices in Bangladesh. The data show consistent intra-class visual characteristics such as leaf shape, venation patterns, texture, margin structure, and colour distribution while exhibiting clear inter-species differences that support reliable species-level classification and comparative analysis. All images were manually collected from natural outdoor environments, including agricultural fields, nurseries, and home gardens located in the Akran Bazar, Ashulia, and Savar regions of Dhaka, Bangladesh. This real-world acquisition approach preserves natural variations in illumination, background complexity, leaf orientation, scale, and growth stages, making the dataset suitable for developing and evaluating robust recognition models under practical conditions. The dataset is organized into two main folders. The Raw Images folder contains original, unaltered images without background removal or enhancement, preserving authentic acquisition conditions. The Pre-processed Images folder includes processed versions of the raw images to facilitate standardized experimentation; preprocessing steps may include resizing, normalization, and noise reduction, as documented. Images in both folders are arranged in class-wise subdirectories corresponding to each plant species. All images are provided in commonly used formats, allowing researchers to apply custom preprocessing, feature extraction, and modelling strategies. The dataset can serve as a benchmark resource for evaluating classification robustness and generalization performance. In addition, it enables comparative studies across different algorithms, supports dataset augmentation research, and facilitates reproducible experimentation. The dataset is suitable for academic research, educational purposes, and applied studies in medicinal plant identification, biodiversity conservation, agricultural technology, and healthcare-oriented artificial intelligence research, particularly in regions where publicly available medicinal plant datasets remain limited.
本数据集收录了经精心筛选整理的孟加拉国本土7种常用药用植物的叶片图像,旨在为基于计算机视觉(Computer Vision)与机器学习(Machine Learning)技术的植物识别与分类研究提供支撑。本研究的核心假设为:在真实户外环境下,叶片形态学特征可提供足够的视觉信息以区分不同药用植物物种。 本数据集共包含13095张叶片图像,其中原始图像6700张、预处理图像6395张,所有图像均被统一调整至640×640像素。此外,本次采集的图像覆盖孟加拉国传统与现代医疗实践中广泛使用的7种药用植物。 该数据集的图像具备稳定的类内视觉特征,包括叶片形状、脉纹样式、纹理、叶缘结构与色彩分布,同时呈现出清晰的种间差异,可为可靠的物种级分类与对比分析提供支持。 所有图像均为手动采集自孟加拉国达卡市阿克兰巴扎、阿舒利亚和萨瓦尔区域的自然户外环境,涵盖农田、苗圃与家庭花园。这种真实场景的采集方式保留了光照、背景复杂度、叶片朝向、尺度与生长阶段等自然变量,使得本数据集适用于在实际应用场景下开发与评估鲁棒性较强的识别模型。 本数据集分为两个主要文件夹:「原始图像(Raw Images)」文件夹包含未经过任何修改的原始图像,未进行背景移除或增强处理,完整保留了采集时的真实环境信息;「预处理图像(Pre-processed Images)」文件夹则包含原始图像的处理后版本,以支持标准化实验,预处理步骤包括图像尺寸调整、归一化与降噪(详见文档说明)。两个文件夹内的图像均按照植物物种划分至对应的类级子目录中。 所有图像均采用通用格式存储,研究人员可根据需求自定义预处理、特征提取与建模策略。 本数据集可作为评估分类模型鲁棒性与泛化性能的基准资源,此外还可用于不同算法间的对比研究、数据集增强相关研究,并支持可复现的实验开展。 本数据集适用于药用植物识别、生物多样性保护、农业技术以及医疗导向人工智能研究中的学术研究、教学用途与应用研究,尤其适用于公开药用植物数据集较为匮乏的地区。



