遇见数据集

SorghumWeedDataset_Segmentation

收藏
Mendeley Data2024-01-31 更新2024-06-26 收录
官方服务:

资源简介:

Purpose of dataset creation: ‘SorghumWeedDataset_Segmentation’ is created to address real-time weed challenges precisely and encourage weed research using computer vision applications. About the dataset: ‘SorghumWeedDataset_Segmentation’ is a crop-weed research dataset with 5555 manually pixel-wise annotated data segments from 252 data samples which can be used for object detection, instance segmentation, and semantic segmentation. The data segments consist of sorghum samplings (Class 0), Grasses (Class 1), and Broad-leaf weeds (Class 2) which are the three research objects focused during this data acquisition process. The TVT (Train: Validate: Test) ratio is set as 8:1:1 to split the data samples into training, validation, and testing. The ground truth preparation is carried out by manually annotating the data segments pixel-wise, using VIA (VGG Image Annotator) software. The respective annotation files for training, validation, and testing are provided in JSON, CSV, and COCO formats. Equipment used for data acquisition: To record a rich set of information on the research objects, a state-of-the-art instrument - Canon EOS 80D – a Digital Single Lens Reflex (DSLR) camera with a sensor type of 22.3mm x 14.9 mm CMOS is used. Data type, format, and size: Each data sample is an RGB image represented in JPEG format with 6000 × 4000 pixels making an average size of 13MB. This rich set of information from the data sample assisted in annotating data segments of plant length lesser than 0.5cm. Temporal coverage: Data is acquired during April and May 2023. To generalize the dataset, data is acquired in various light and weather conditions with varying distances. Geographical coverage: Data is acquired from Sri Ramaswamy Memorial (SRM) Care Farm, Chengalpattu district, Tamil Nadu, India. To the best of our knowledge, ‘SorghumWeedDataset_Segmentation’ is the first open-access crop-weed research dataset from Indian fields for segmentation that deals with weed issues in uniform and random crop-spacing fields. Expected outcome: The expected outcome of this dataset will be an Artificial Intelligence (AI) model that localizes and segments all the research objects present in a particular data sample. Detailed description: A detailed description of the dataset and data acquisition process is given in the data article entitled “ ‘SorghumWeedDataset_Classification’ And ‘SorghumWeedDataset_Segmentation’ Datasets For Classification, Detection, and Segmentation In Deep Learning “. (Submitted in the journal ‘Data in Brief’ on 25/09/2023 and awaiting publication) Citation: If you find this dataset helpful and use it in your work, kindly cite this dataset using “Michael, Justina; M, Thenmozhi (2023), “SorghumWeedDataset_Segmentation”, Mendeley Data, V1, doi: 10.17632/y9bmtf4xmr.1” Further queries: If any queries/suggestions concerning this dataset, please e-mail us at thenmozm@srmist.edu.in [corresponding author]

数据集创建目的:‘SorghumWeedDataset_Segmentation’数据集的创建目标为精准应对实时杂草治理难题,推动基于计算机视觉应用的杂草相关研究。 数据集概况:‘SorghumWeedDataset_Segmentation’是一款作物-杂草研究专用数据集,源自252组原始数据样本,包含5555份经人工逐像素标注的数据片段,可应用于目标检测、实例分割与语义分割任务。数据片段涵盖三类核心研究对象:高粱植株(类别0)、禾本科杂草(类别1)与阔叶杂草(类别2),为本轮数据采集阶段的聚焦目标。数据集采用8:1:1的训练集-验证集-测试集(TVT)划分比例,将原始样本拆分为训练、验证与测试子集。标注真值(ground truth)通过人工逐像素标注制作,使用VGG图像标注器(VGG Image Annotator,VIA)软件完成。训练、验证与测试集对应的标注文件分别以JSON、CSV与COCO格式提供。 数据采集设备:为完整记录研究对象的丰富信息,本次采集采用当前先进的成像设备——佳能EOS 80D数码单反相机(Digital Single Lens Reflex,DSLR),其传感器尺寸为22.3mm×14.9mm,采用CMOS传感器类型。 数据类型、格式与大小:每份原始数据样本为RGB图像,以JPEG格式存储,分辨率为6000×4000像素,平均单文件大小约13MB。该高分辨率数据集支持对长度小于0.5cm的植株片段进行精准标注。 时间覆盖范围:数据采集于2023年4月至5月期间。为提升数据集的泛化能力,采集过程涵盖了不同光照、天气条件与拍摄距离的多样化场景。 地理覆盖范围:数据采集自印度泰米尔纳德邦琴格伯杜(Chengalpattu)地区的斯里·拉马斯瓦米纪念(SRM)护养农场。据我们所知,‘SorghumWeedDataset_Segmentation’是首个源自印度农田、面向分割任务的开放获取作物-杂草研究数据集,可用于解决均匀与随机作物行距田间的杂草治理问题。 预期成果:本数据集的预期应用成果为一款人工智能(Artificial Intelligence,AI)模型,可实现对特定数据样本中所有研究对象的精准定位与分割。 详细说明:本数据集与数据采集流程的详细描述,请参阅已提交至《Data in Brief》期刊的论文《‘SorghumWeedDataset_Classification’与‘SorghumWeedDataset_Segmentation’数据集:深度学习中的分类、检测与分割任务》(2023年9月25日提交,待刊)。 引用说明:若您认为本数据集对研究有所帮助并将其应用于相关工作,请使用以下信息进行引用:Michael, Justina; M, Thenmozhi (2023), "SorghumWeedDataset_Segmentation", Mendeley Data, V1, doi: 10.17632/y9bmtf4xmr.1 后续咨询:若您对本数据集有任何疑问或建议,请通过通讯作者邮箱thenmozm@srmist.edu.in与我们联系。

创建时间:
2024-01-31
搜集汇总
数据集介绍
SorghumWeedDataset_Segmentation 数据集图片
背景与挑战
背景概述
SorghumWeedDataset_Segmentation是一个用于计算机视觉作物-杂草研究的开放存取数据集,包含5555个手动像素级标注的数据片段,涵盖高粱、禾本科杂草和阔叶杂草三类对象,适用于目标检测、实例分割和语义分割任务。数据集采集于2023年4月至5月的印度田间,使用高分辨率相机拍摄,图像质量高,并按8:1:1比例划分为训练、验证和测试集,提供多种标注格式。该数据集是首个针对印度田间杂草问题的开放存取分割数据集,旨在支持AI模型开发以精准定位和分割杂草对象。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务