遇见数据集

SorghumWeedDataset_Segmentation

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Purpose of dataset creation: ‘SorghumWeedDataset_Segmentation’ is created to address real-time weed challenges precisely and encourage weed research using computer vision applications. About the dataset: ‘SorghumWeedDataset_Segmentation’ is a crop-weed research dataset with 5555 manually pixel-wise annotated data segments from 252 data samples which can be used for object detection, instance segmentation, and semantic segmentation. The data segments consist of sorghum samplings (Class 0), Grasses (Class 1), and Broad-leaf weeds (Class 2) which are the three research objects focused during this data acquisition process. The TVT (Train: Validate: Test) ratio is set as 8:1:1 to split the data samples into training, validation, and testing. The ground truth preparation is carried out by manually annotating the data segments pixel-wise, using VIA (VGG Image Annotator) software. The respective annotation files for training, validation, and testing are provided in JSON, CSV, and COCO formats. Equipment used for data acquisition: To record a rich set of information on the research objects, a state-of-the-art instrument - Canon EOS 80D – a Digital Single Lens Reflex (DSLR) camera with a sensor type of 22.3mm x 14.9 mm CMOS is used. Data type, format, and size: Each data sample is an RGB image represented in JPEG format with 6000 × 4000 pixels making an average size of 13MB. This rich set of information from the data sample assisted in annotating data segments of plant length lesser than 0.5cm. Temporal coverage: Data is acquired during April and May 2023. To generalize the dataset, data is acquired in various light and weather conditions with varying distances. Geographical coverage: Data is acquired from Sri Ramaswamy Memorial (SRM) Care Farm, Chengalpattu district, Tamil Nadu, India. To the best of our knowledge, ‘SorghumWeedDataset_Segmentation’ is the first open-access crop-weed research dataset from Indian fields for segmentation that deals with weed issues in uniform and random crop-spacing fields. Expected outcome: The expected outcome of this dataset will be an Artificial Intelligence (AI) model that localizes and segments all the research objects present in a particular data sample. Detailed description: A detailed description of the dataset and data acquisition process is given in the data article entitled “ ‘SorghumWeedDataset_Classification’ And ‘SorghumWeedDataset_Segmentation’ Datasets For Classification, Detection, and Segmentation In Deep Learning “. (Submitted in the journal ‘Data in Brief’ on 25/09/2023 and awaiting publication) Citation: If you find this dataset helpful and use it in your work, kindly cite this dataset using “Michael, Justina; M, Thenmozhi (2023), “SorghumWeedDataset_Segmentation”, Mendeley Data, V1, doi: 10.17632/y9bmtf4xmr.1” Further queries: If any queries/suggestions concerning this dataset, please e-mail us at thenmozm@srmist.edu.in [corresponding author]

数据集创建目的: ‘SorghumWeedDataset_Segmentation’ 数据集旨在精准解决实时杂草防控难题,推动计算机视觉技术在杂草相关研究中的应用。 数据集概况: ‘SorghumWeedDataset_Segmentation’ 是一款作物-杂草研究数据集,包含来自252份原始数据样本的5555份人工像素级标注数据片段,可用于目标检测、实例分割与语义分割任务。本次数据采集聚焦三类研究对象:高粱植株(类别0)、禾本科杂草(类别1)与阔叶杂草(类别2)。数据集采用训练集:验证集:测试集(Train: Validate: Test, TVT)=8:1:1的比例,将数据样本划分为训练、验证与测试子集。标注文件通过对数据片段进行像素级手动标注生成,所用工具为VGG图像标注器(VGG Image Annotator, VIA)。训练、验证及测试集对应的标注文件分别以JSON、CSV与COCO格式提供。 数据采集设备: 本研究采用当前主流的佳能EOS 80D数码单反(Digital Single Lens Reflex, DSLR)相机进行数据采集,该相机搭载尺寸为22.3mm×14.9mm的CMOS传感器。 数据类型、格式与尺寸: 每份原始数据样本为RGB图像,采用JPEG格式存储,分辨率为6000×4000像素,单张图像平均大小约为13MB。该数据集包含的丰富信息可支撑对株高不足0.5cm的植株进行数据片段标注。 时间覆盖范围: 数据采集时间为2023年4月至5月。为提升数据集的泛化能力,采集过程涵盖了不同光照、天气条件以及不同拍摄距离的场景。 地理覆盖范围: 数据采集地点位于印度泰米尔纳德邦陈加尔帕杜(Chengalpattu)区的斯里拉姆斯瓦米纪念(Sri Ramaswamy Memorial, SRM)护理农场。据我们所知,‘SorghumWeedDataset_Segmentation’ 是首个源自印度农田、面向分割任务的开放获取作物-杂草研究数据集,可用于解决均匀与随机种植间距农田中的杂草问题。 预期应用成果: 本数据集的预期应用成果为一款人工智能(Artificial Intelligence, AI)模型,可实现对单份数据样本中所有研究对象的定位与分割。 详细说明: 本数据集与数据采集流程的详细说明已发表于题为《‘SorghumWeedDataset_Classification’与‘SorghumWeedDataset_Segmentation’数据集:深度学习中的分类、检测与分割任务数据集》的数据类文章(2023年9月25日提交至《Data in Brief》期刊,待刊)。 引用规范: 若您在研究工作中使用本数据集,请按以下方式引用:Michael, Justina; M, Thenmozhi (2023), "SorghumWeedDataset_Segmentation", Mendeley Data, V1, doi: 10.17632/y9bmtf4xmr.1 咨询方式: 若您对本数据集有任何疑问或建议,请联系通讯作者发送邮件至thenmozm@srmist.edu.in。

创建时间:
2023-09-26
二维码
社区交流群
二维码
科研交流群
商业服务