遇见数据集

Synthetic Meets Authentic: Leveraging Text-to-Image Generated Datasets for Apple Detection in Orchard Environments

收藏
Mendeley Data2024-03-28 更新2024-06-29 收录
官方服务:

资源简介:

Training machine learning (ML) models for computer vision-based object detection process typically requires large, labeled datasets, a process often burdened by significant human effort and high costs associated with imaging systems and image acquisition. This research aimed to simplify image data collection for object detection in orchards by avoiding traditional fieldwork with different imaging sensors. Utilizing OpenAI's DALLE, a large language model (LLM) for realistic image generation, we generated and annotated a cost effective dataset. This dataset, exclusively generated with text-to-image prompts/inputs, was then utilized to train a deep learning model, YOLOv8, for apple detection, which was then tested with real-world (outdoor orchard) images captured by a digital (Nikon D5100) camera as well as a machine vision camera (IntelRealsense D435i). The model achieved a training precision of 0.83, recall of 0.99, an F1 score of 0.92, and mAP@50 at 0.96. Validation tests against actual images collected over two different varieties of apples (Honeycrisp and Envy) in a commercial orchard environment showed a precision of 0.82 and 0.75, recall of 0.88 and 0.63, and mAP@50 of 0.92 and 0.70, each respectively. The inference time of the model was 0.015 seconds for the digital camera-based images and 0.012 seconds for the machine vision camera based images. This study presents a pathway for generating large image datasets in challenging agricultural fields with minimal or no labor-intensive efforts in field data-collection, which could accelerate the development and deployment of computer vision and robotic technologies in orchard environments.

针对基于计算机视觉的目标检测任务训练机器学习(ML, Machine Learning)模型,通常需要大规模带标注数据集,而该流程往往伴随大量人力投入,以及成像系统与图像采集相关的高额成本。本研究旨在简化果园目标检测的图像数据采集流程,无需依赖传统野外作业及各类成像传感器。本研究借助OpenAI的DALLE——一款用于生成写实图像的大语言模型(LLM, Large Language Model)——生成并标注了一套高性价比数据集。该数据集完全通过文本到图像提示词/输入生成,随后被用于训练用于苹果检测的深度学习模型YOLOv8。模型的测试环节采用了由数码单反相机(尼康D5100, Nikon D5100)与机器视觉相机(英特尔实感D435i, IntelRealSense D435i)采集的真实户外果园图像。该模型在训练阶段的精确率达0.83、召回率0.99、F1分数0.92,且mAP@50值为0.96。针对商业果园环境下采集的两种苹果品种(Honeycrisp, 蜜脆;Envy, 爱妃)的真实图像开展的验证测试结果显示,模型的精确率分别为0.82与0.75,召回率分别为0.88与0.63,mAP@50值分别为0.92与0.70。针对数码单反相机采集图像的模型推理时长为0.015秒,针对机器视觉相机采集图像的推理时长则为0.012秒。本研究为在复杂农田场景中通过极低甚至无需野外数据采集的人力密集型投入来生成大规模图像数据集提供了可行路径,有望加速计算机视觉与机器人技术在果园环境中的研发与部署。

创建时间:
2024-03-27
二维码
社区交流群
二维码
科研交流群
商业服务