simonMadec/VegAnn
收藏资源简介:
--- language: - en size_categories: - 1K<n<10K task_categories: - image-segmentation tags: - vegetation - segmentation DOI: - 10.1038/s41597-023-02098-y licence: - CC-BY dataset_info: features: - name: image dtype: image - name: mask dtype: image - name: System dtype: string - name: Orientation dtype: string - name: latitude dtype: float64 - name: longitude dtype: float64 - name: date dtype: string - name: LocAcc dtype: int64 - name: Species dtype: string - name: Owner dtype: string - name: Dataset-Name dtype: string - name: TVT-split1 dtype: string - name: TVT-split2 dtype: string - name: TVT-split3 dtype: string - name: TVT-split4 dtype: string - name: TVT-split5 dtype: string splits: - name: train num_bytes: 1896819757.9 num_examples: 3775 download_size: 1940313757 dataset_size: 1896819757.9 configs: - config_name: default data_files: - split: train path: data/train-* --- # VegAnn Dataset ### **Vegetation Annotation of a large multi-crop RGB Dataset acquired under diverse conditions for image semantic segmentation** ## Keypoints ⏳ - VegAnn contains 3775 images - Images are 512*512 pixels - Corresponding binary masks is 0 for soil + crop residues (background) 255 for Vegetation (foreground) - The dataset includes images of 26+ crop species, which are not evenly represented - VegAnn was compiled using a variety of outdoor images captured with different acquisition systems and configurations - For more information about VegAnn, details, labeling rules and potential uses see https://doi.org/10.1038/s41597-023-02098-y ## Dataset Description 📚 VegAnn, short for Vegetation Annotation, is a meticulously curated collection of 3,775 multi-crop RGB images aimed at enhancing research in crop vegetation segmentation. These images span various phenological stages and were captured using diverse systems and platforms under a wide range of illumination conditions. By aggregating sub-datasets from different projects and institutions, VegAnn represents a broad spectrum of measurement conditions, crop species, and development stages. ### Languages 🌐 The annotations and documentation are primarily in English. ## Dataset Structure 🏗 ### Data Instances 📸 A VegAnn data instance consists of a 512x512 pixel RGB image patch derived from larger raw images. These patches are designed to provide sufficient detail for distinguishing between vegetation and background, crucial for applications in semantic segmentation and other forms of computer vision analysis in agricultural contexts.  ### Data Fields 📋 - `Name`: Unique identifier for each image patch. - `System`: The imaging system used to acquire the photo (e.g., Handheld Cameras, DHP, UAV). - `Orientation`: The camera's orientation during image capture (e.g., Nadir, 45 degrees). - `latitude` and `longitude`: Geographic coordinates where the image was taken. - `date`: Date of image acquisition. - `LocAcc`: Location accuracy flag (1 for high accuracy, 0 for low or uncertain accuracy). - `Species`: The crop species featured in the image (e.g., Wheat, Maize, Soybean). - `Owner`: The institution or entity that provided the image (e.g., Arvalis, INRAe). - `Dataset-Name`: The sub-dataset or project from which the image originates (e.g., Phenomobile, Easypcc). - `TVT-split1` to `TVT-split5`: Fields indicating the train/validation/test split configurations, facilitating various experimental setups. ### Data Splits 📊 The dataset is structured into multiple splits (as indicated by `TVT-split` fields) to support different training, validation, and testing scenarios in machine learning workflows. ## Dataset Creation 🛠 ### Curation Rationale 🤔 The VegAnn dataset was developed to address the gap in available datasets for training convolutional neural networks (CNNs) for the task of semantic segmentation in real-world agricultural environments. By incorporating images from a wide array of conditions and stages of crop development, VegAnn aims to enhance the performance of segmentation algorithms, promote benchmarking, and foster research on large-scale crop vegetation segmentation. ### Source Data 🌱 #### Initial Data Collection and Normalization Images within VegAnn were sourced from various sub-datasets contributed by different institutions, each under specific acquisition configurations. These were then standardized into 512x512 pixel patches to maintain consistency across the dataset. #### Who are the source data providers? The data was provided by a collaboration of institutions including Arvalis, INRAe, The University of Tokyo, University of Queensland, NEON, and EOLAB, among others.  ### Annotations 📝 #### Annotation process Annotations for the dataset were focused on distinguishing between vegetation and background within the images. The process ensured that the images offered sufficient spatial resolution to allow for accurate visual segmentation. #### Who are the annotators? The annotations were performed by a team comprising researchers and domain experts from the contributing institutions. ## Considerations for Using the Data 🤓 ### Social Impact of Dataset 🌍 The VegAnn dataset is expected to significantly impact agricultural research and commercial applications by enhancing the accuracy of crop monitoring, disease detection, and yield estimation through improved vegetation segmentation techniques. ### Discussion of Biases 🧐 Given the diverse sources of the images, there may be inherent biases towards certain crop types, geographical locations, and imaging conditions. Users should consider this diversity in applications and analyses. ### Licensing Information 📄 Please refer to the specific licensing agreements of the contributing institutions or contact the dataset providers for more information on usage rights and restrictions. ## Citation Information 📚 If you use the VegAnn dataset in your research, please cite the following: ``` @article{madec_vegann_2023, title = {{VegAnn}, {Vegetation} {Annotation} of multi-crop {RGB} images acquired under diverse conditions for segmentation}, volume = {10}, issn = {2052-4463}, url = {https://doi.org/10.1038/s41597-023-02098-y}, doi = {10.1038/s41597-023-02098-y}, abstract = {Applying deep learning to images of cropping systems provides new knowledge and insights in research and commercial applications. Semantic segmentation or pixel-wise classification, of RGB images acquired at the ground level, into vegetation and background is a critical step in the estimation of several canopy traits. Current state of the art methodologies based on convolutional neural networks (CNNs) are trained on datasets acquired under controlled or indoor environments. These models are unable to generalize to real-world images and hence need to be fine-tuned using new labelled datasets. This motivated the creation of the VegAnn - Vegetation Annotation - dataset, a collection of 3775 multi-crop RGB images acquired for different phenological stages using different systems and platforms in diverse illumination conditions. We anticipate that VegAnn will help improving segmentation algorithm performances, facilitate benchmarking and promote large-scale crop vegetation segmentation research.}, number = {1}, journal = {Scientific Data}, author = {Madec, Simon and Irfan, Kamran and Velumani, Kaaviya and Baret, Frederic and David, Etienne and Daubige, Gaetan and Samatan, Lucas Bernigaud and Serouart, Mario and Smith, Daniel and James, Chrisbin and Camacho, Fernando and Guo, Wei and De Solan, Benoit and Chapman, Scott C. and Weiss, Marie}, month = may, year = {2023}, pages = {302}, } ``` ## Additional Information - **Dataset Curators**: Simon Madec et al. - **Version**: 1.0 - **License**: CC-BY - **Contact**: simon.madec@cirad.fr
VegAnn 数据集概述
数据集描述
VegAnn(Vegetation Annotation)是一个精心策划的多作物RGB图像集合,旨在促进作物植被分割的研究。该数据集包含3,775张图像,涵盖了各种物候阶段,并使用多种系统和平台在广泛的照明条件下捕获。通过汇集来自不同项目和机构的子数据集,VegAnn代表了广泛的测量条件、作物种类和发展阶段。
关键点
- 包含3,775张图像
- 图像分辨率为512x512像素
- 对应的二值掩码:0表示土壤和作物残留物(背景),255表示植被(前景)
- 数据集包括26种以上的作物,分布不均
- 图像由不同的采集系统和配置在户外捕获
数据结构
数据实例
每个VegAnn数据实例由一个512x512像素的RGB图像块组成,这些图像块是从更大的原始图像中提取的,旨在提供足够的细节以区分植被和背景,这对于农业环境中的语义分割和其他计算机视觉分析至关重要。
数据字段
Name:每个图像块的唯一标识符System:用于获取照片的成像系统(例如,手持相机、DHP、UAV)Orientation:图像捕获时相机的方向(例如,正下方、45度)latitude和longitude:图像拍摄的地理坐标date:图像获取日期LocAcc:位置精度标志(1表示高精度,0表示低或不确定精度)Species:图像中的作物种类(例如,小麦、玉米、大豆)Owner:提供图像的机构或实体(例如,Arvalis、INRAe)Dataset-Name:图像来源的子数据集或项目(例如,Phenomobile、Easypcc)TVT-split1到TVT-split5:指示训练/验证/测试分割配置的字段,便于各种实验设置
数据分割
数据集被划分为多个分割(由TVT-split字段指示),以支持机器学习工作流中的不同训练、验证和测试场景。
数据集创建
策划理由
VegAnn数据集的开发旨在填补现有数据集在训练卷积神经网络(CNNs)进行真实农业环境中的语义分割任务方面的空白。通过包含来自广泛条件和作物发展阶段的图像,VegAnn旨在提高分割算法的性能,促进基准测试,并推动大规模作物植被分割研究。
源数据
初始数据收集和标准化
VegAnn中的图像来自不同机构贡献的多个子数据集,每个子数据集在特定的采集配置下捕获。这些图像随后被标准化为512x512像素的图像块,以保持数据集的一致性。
源数据提供者
数据由包括Arvalis、INRAe、东京大学、昆士兰大学、NEON和EOLAB等机构的合作提供。
标注
标注过程
数据集的标注集中在区分图像中的植被和背景。标注过程确保图像提供足够的空间分辨率,以便进行准确的视觉分割。
标注者
标注工作由来自贡献机构的研究人员和领域专家组成的团队执行。
使用数据集的注意事项
数据集的社会影响
VegAnn数据集预计将显著影响农业研究和商业应用,通过改进植被分割技术提高作物监测、疾病检测和产量估计的准确性。
讨论偏差
鉴于图像来源的多样性,可能存在对某些作物类型、地理位置和成像条件的固有偏差。用户在应用和分析时应考虑这种多样性。
许可信息
请参考贡献机构的特定许可协议,或联系数据集提供者以获取更多关于使用权利和限制的信息。
引用信息
如果您在研究中使用VegAnn数据集,请引用以下内容:
@article{madec_vegann_2023, title = {{VegAnn}, {Vegetation} {Annotation} of multi-crop {RGB} images acquired under diverse conditions for segmentation}, volume = {10}, issn = {2052-4463}, url = {https://doi.org/10.1038/s41597-023-02098-y}, doi = {10.1038/s41597-023-02098-y}, abstract = {Applying deep learning to images of cropping systems provides new knowledge and insights in research and commercial applications. Semantic segmentation or pixel-wise classification, of RGB images acquired at the ground level, into vegetation and background is a critical step in the estimation of several canopy traits. Current state of the art methodologies based on convolutional neural networks (CNNs) are trained on datasets acquired under controlled or indoor environments. These models are unable to generalize to real-world images and hence need to be fine-tuned using new labelled datasets. This motivated the creation of the VegAnn - Vegetation Annotation - dataset, a collection of 3775 multi-crop RGB images acquired for different phenological stages using different systems and platforms in diverse illumination conditions. We anticipate that VegAnn will help improving segmentation algorithm performances, facilitate benchmarking and promote large-scale crop vegetation segmentation research.}, number = {1}, journal = {Scientific Data}, author = {Madec, Simon and Irfan, Kamran and Velumani, Kaaviya and Baret, Frederic and David, Etienne and Daubige, Gaetan and Samatan, Lucas Bernigaud and Serouart, Mario and Smith, Daniel and James, Chrisbin and Camacho, Fernando and Guo, Wei and De Solan, Benoit and Chapman, Scott C. and Weiss, Marie}, month = may, year = {2023}, pages = {302}, }
附加信息
- 数据集策展人:Simon Madec 等人
- 版本:1.0
- 许可证:CC-BY
- 联系:simon.madec@cirad.fr




