遇见数据集

avkhalkar/PlantVillage

收藏
Hugging Face2026-04-04 更新2026-04-12 收录
官方服务:

资源简介:

--- language: - en license: cc-by-sa-3.0 task_categories: - image-classification tags: - agriculture - plant-disease - biology dataset_info: features: - name: image dtype: image - name: image_path dtype: string - name: label dtype: class_label: names: '0': Apple___Apple_scab '1': Apple___Black_rot '2': Apple___Cedar_apple_rust '3': Apple___healthy '4': Blueberry___healthy '5': Cherry_(including_sour)___Powdery_mildew '6': Cherry_(including_sour)___healthy '7': Corn_(maize)___Cercospora_leaf_spot Gray_leaf_spot '8': Corn_(maize)___Common_rust_ '9': Corn_(maize)___Northern_Leaf_Blight '10': Corn_(maize)___healthy '11': Grape___Black_rot '12': Grape___Esca_(Black_Measles) '13': Grape___Leaf_blight_(Isariopsis_Leaf_Spot) '14': Grape___healthy '15': Orange___Haunglongbing_(Citrus_greening) '16': Peach___Bacterial_spot '17': Peach___healthy '18': Pepper,_bell___Bacterial_spot '19': Pepper,_bell___healthy '20': Potato___Early_blight '21': Potato___Late_blight '22': Potato___healthy '23': Raspberry___healthy '24': Soybean___healthy '25': Squash___Powdery_mildew '26': Strawberry___Leaf_scorch '27': Strawberry___healthy '28': Tomato___Bacterial_spot '29': Tomato___Early_blight '30': Tomato___Late_blight '31': Tomato___Leaf_Mold '32': Tomato___Septoria_leaf_spot '33': Tomato___Spider_mites Two-spotted_spider_mite '34': Tomato___Target_Spot '35': Tomato___Tomato_Yellow_Leaf_Curl_Virus '36': Tomato___Tomato_mosaic_virus '37': Tomato___healthy - name: crop dtype: string - name: disease dtype: string - name: leaf_id dtype: string config_name: color splits: - name: train num_bytes: 43596 num_examples: 43596 - name: test num_bytes: 10709 num_examples: 10709 download_size: 2000000000 dataset_size: 2000000000 --- # PlantVillage Dataset [![Paper](https://img.shields.io/badge/Paper-Read-green)](https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2016.01419/full) [![GitHub](https://img.shields.io/badge/GitHub-Repository-black)](https://github.com/spMohanty/PlantVillage-Dataset) <img src="https://raw.githubusercontent.com/spMohanty/PlantVillage-Dataset/master/generated_for_paper/plantvillage.jpg" alt="PlantVillage Dataset Sample" width="600"/> The **PlantVillage Dataset** is an open access repository of **54,306 images** of healthy and diseased plant leaves, collected to advance research in automated plant disease diagnosis. It covers **14 crop species** and **26 diseases**. This dataset was introduced in the paper [**"Using Deep Learning for Image-Based Plant Disease Detection"**](https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2016.01419/full) by Mohanty et al. (2016). ## Quick Start The dataset comes with pre-defined **80/20 train/test splits** that preserve the leaf grouping logic (ensuring images of the same leaf do not appear in both sets). ```python from datasets import load_dataset # Load the default configuration (color images) # This automatically downloads the train and test splits. dataset = load_dataset("mohanty/PlantVillage", "color") print(dataset) # DatasetDict({ # train: Dataset({ features: [...], num_rows: 43596 }), # test: Dataset({ features: [...], num_rows: 10709 }) # }) ``` ## Dataset Configurations You can choose from three configurations depending on your needs: | Configuration | Description | Usage | |---|---|---| | **`color`** | Original RGB images (Default) | `load_dataset("mohanty/PlantVillage", "color")` | | **`grayscale`** | Grayscale versions | `load_dataset("mohanty/PlantVillage", "grayscale")` | | **`segmented`** | Background removed, leaf segmented | `load_dataset("mohanty/PlantVillage", "segmented")` | ## Advanced: Custom Splitting **Note:** The dataset **already includes** a standard train/test split (as shown in Quick Start), which is recommended for benchmarking. The instructions below are only for **advanced users** who require custom cross-validation folds. If you require a different split ratio or cross-validation scheme, you **must** strictly respect the `leaf_id` to prevent data leakage. Multiple images often capture the same physical leaf; separating them across train/test sets will bias your evaluation. Here is a complete example of how to reshuffle and split the dataset 80/20: ```python import numpy as np from datasets import load_dataset, concatenate_datasets # 1. Load the full dataset (combining default train/test splits) dataset = load_dataset("mohanty/PlantVillage", "color") full_dataset = concatenate_datasets([dataset["train"], dataset["test"]]) # 2. Get unique leaf IDs representing physical leaves all_leaf_ids = np.unique(full_dataset["leaf_id"]) # 3. Shuffle and Split Leaf IDs (e.g., 80% train, 20% test) np.random.seed(42) np.random.shuffle(all_leaf_ids) split_ratio = 0.8 split_idx = int(len(all_leaf_ids) * split_ratio) train_leaf_ids = set(all_leaf_ids[:split_idx]) test_leaf_ids = set(all_leaf_ids[split_idx:]) # 4. Filter the full dataset to create new splits # This ensures all images of a specific leaf are exclusively in one split custom_train = full_dataset.filter(lambda x: x["leaf_id"] in train_leaf_ids) custom_test = full_dataset.filter(lambda x: x["leaf_id"] in test_leaf_ids) print(f"Custom Train Size: {len(custom_train)}") print(f"Custom Test Size: {len(custom_test)}") ``` ## Dataset Features - **`image`**: PIL Image. - **`label`**: Class label (e.g., `Apple___Black_rot`). - **`leaf_id`**: Unique identifier for the physical leaf. - **`crop`**: Crop name. - **`disease`**: Disease name. ## Citation ```bibtex @article{Mohanty_Hughes_Salathé_2016, title = {Using deep learning for image-based plant disease detection}, volume = {7}, DOI = {10.3389/fpls.2016.01419}, journal = {Frontiers in Plant Science}, author = {Mohanty, Sharada P. and Hughes, David P. and Salathé, Marcel}, year = {2016}, month = {Sep} } ``` ## Author Sharada Mohanty <sharada.mohanty@epfl.ch> Marcel Salathé <Marcel.Salathe@epfl.ch> **Digital Epidemiology Lab, EPFL**

--- 语言: - en(英语) 许可证:CC BY-SA 3.0(知识共享署名-相同方式共享3.0协议) 任务类别: - image-classification(图像分类) 标签: - 农业 - 植物病害 - 生物学 dataset_info: 特征: - 名称:image 数据类型:image(图像) - 名称:image_path 数据类型:string(字符串) - 名称:label 数据类型: class_label(类别标签): 名称映射: '0': 苹果_苹果黑星病 '1': 苹果_苹果腐烂病 '2': 苹果_雪松苹果锈病 '3': 苹果_健康植株 '4': 蓝莓_健康植株 '5': 樱桃(含酸樱桃)_白粉病 '6': 樱桃(含酸樱桃)_健康植株 '7': 玉米_尾孢叶斑病(灰斑病) '8': 玉米_普通锈病 '9': 玉米_北方叶枯病 '10': 玉米_健康植株 '11': 葡萄_黑腐病 '12': 葡萄_埃斯卡病(黑麻疹病) '13': 葡萄_叶枯病(柱孢叶斑病) '14': 葡萄_健康植株 '15': 橙_柑橘黄龙病(Citrus greening) '16': 桃_细菌性斑点病 '17': 桃_健康植株 '18': 甜椒_细菌性斑点病 '19': 甜椒_健康植株 '20': 马铃薯_早疫病 '21': 马铃薯_晚疫病 '22': 马铃薯_健康植株 '23': 树莓_健康植株 '24': 大豆_健康植株 '25': 南瓜_白粉病 '26': 草莓_叶灼病 '27': 草莓_健康植株 '28': 番茄_细菌性斑点病 '29': 番茄_早疫病 '30': 番茄_晚疫病 '31': 番茄_叶霉病 '32': 番茄_壳针孢叶斑病 '33': 番茄_二斑叶螨危害 '34': 番茄_靶斑病 '35': 番茄_番茄黄化曲叶病毒病 '36': 番茄_番茄花叶病毒病 '37': 番茄_健康植株 - 名称:crop 数据类型:string(字符串) - 名称:disease 数据类型:string(字符串) - 名称:leaf_id 数据类型:string(字符串) 配置名称:color(彩色配置) 划分集: - 名称:train(训练集) 字节数:43596 样本数:43596 - 名称:test(测试集) 字节数:10709 样本数:10709 下载大小:2000000000字节 数据集总大小:2000000000字节 --- # PlantVillage 数据集(PlantVillage Dataset) [![论文](https://img.shields.io/badge/Paper-阅读-green)](https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2016.01419/full) [![GitHub仓库](https://img.shields.io/badge/GitHub-仓库-black)](https://github.com/spMohanty/PlantVillage-Dataset) <img src="https://raw.githubusercontent.com/spMohanty/PlantVillage-Dataset/master/generated_for_paper/plantvillage.jpg" alt="PlantVillage 数据集示例" width="600"/> **PlantVillage 数据集(PlantVillage Dataset)** 是一个开放获取的资源库,包含54306张健康与染病植物叶片图像,旨在推动自动化植物病害诊断领域的研究。该数据集涵盖14种作物与26种病害。 本数据集由Mohanty等人于2016年发表的论文《**基于深度学习的图像式植物病害检测**》(原文标题:Using Deep Learning for Image-Based Plant Disease Detection)中首次提出。 ## 快速上手 本数据集已预设80/20的训练集/测试集划分规则,该规则保留叶片分组逻辑(确保同一叶片的图像不会同时出现在训练集与测试集中)。 python from datasets import load_dataset # 加载默认配置(彩色图像) # 该命令将自动下载训练集与测试集 dataset = load_dataset("mohanty/PlantVillage", "color") print(dataset) # DatasetDict({ # train: Dataset({ features: [...], num_rows: 43596 }), # test: Dataset({ features: [...], num_rows: 10709 }) # }) ## 数据集配置 您可根据需求选择三种配置方案: | 配置名称 | 描述 | 使用方法 | |---|---|---| | **`color`** | 原始RGB彩色图像(默认配置) | `load_dataset("mohanty/PlantVillage", "color")` | | **`grayscale`** | 灰度图像版本 | `load_dataset("mohanty/PlantVillage", "grayscale")` | | **`segmented`** | 移除背景、仅保留叶片的分割图像 | `load_dataset("mohanty/PlantVillage", "segmented")` | ## 进阶:自定义划分 **注意:** 本数据集已包含标准的训练集/测试集划分(如快速上手部分所示),该划分适用于基准测试,推荐优先使用。以下说明仅适用于需要自定义交叉验证折的高级用户。 若您需要不同的划分比例或交叉验证方案,必须严格遵循`leaf_id`(叶片唯一标识)字段进行划分,以避免数据泄露。同一片物理叶片通常会有多张拍摄图像,若将其拆分至不同集合会导致评估结果产生偏差。 以下为完整示例,展示如何将数据集按80/20比例重新洗牌并划分: python import numpy as np from datasets import load_dataset, concatenate_datasets # 1. 加载完整数据集(合并默认的训练集与测试集) dataset = load_dataset("mohanty/PlantVillage", "color") full_dataset = concatenate_datasets([dataset["train"], dataset["test"]]) # 2. 获取代表物理叶片的唯一叶片ID集合 all_leaf_ids = np.unique(full_dataset["leaf_id"]) # 3. 洗牌并划分叶片ID(例如80%用于训练,20%用于测试) np.random.seed(42) np.random.shuffle(all_leaf_ids) split_ratio = 0.8 split_idx = int(len(all_leaf_ids) * split_ratio) train_leaf_ids = set(all_leaf_ids[:split_idx]) test_leaf_ids = set(all_leaf_ids[split_idx:]) # 4. 过滤完整数据集以创建新的划分 # 该操作可确保某一叶片的所有图像均仅属于某一个划分集合 custom_train = full_dataset.filter(lambda x: x["leaf_id"] in train_leaf_ids) custom_test = full_dataset.filter(lambda x: x["leaf_id"] in test_leaf_ids) print(f"自定义训练集大小: {len(custom_train)}") print(f"自定义测试集大小: {len(custom_test)}") ## 数据集字段说明 - **`image`**:PIL格式图像(PIL Image)。 - **`label`**:类别标签(例如`Apple___Black_rot`,即苹果_苹果腐烂病)。 - **`leaf_id`**:物理叶片的唯一标识符。 - **`crop`**:作物名称。 - **`disease`**:病害名称。 ## 引用格式 bibtex @article{Mohanty_Hughes_Salathé_2016, title = {基于深度学习的图像式植物病害检测}, volume = {7}, DOI = {10.3389/fpls.2016.01419}, journal = {Frontiers in Plant Science(植物科学前沿)}, author = {Mohanty, Sharada P. and Hughes, David P. and Salathé, Marcel}, year = {2016}, month = {Sep} } ## 作者 Sharada Mohanty <sharada.mohanty@epfl.ch> Marcel Salathé <Marcel.Salathe@epfl.ch> **瑞士联邦理工学院洛桑分校(EPFL)数字流行病学实验室**

提供机构:
avkhalkar
二维码
社区交流群
二维码
科研交流群
商业服务