Multimodal Agricultural Agent Dataset
收藏资源简介:
本研究构建了一个多模态农业智能体数据集,包含五大任务:分类、检测、视觉问题回答(VQA)、工具选择和智能体评估。数据集涵盖了18种不同的甘蔗病害类别,分类任务有大约10万张注释图像,检测任务有6万多张标签样本。此外,还构建了一个包含超过16万样本的双语(中文和英文)视觉问题回答数据集。该数据集为训练和评估农业视觉语言模型提供了宝贵的资源,并展示了自动生成大规模领域特定VQA数据的有效性。
This study constructs a multimodal agricultural AI agent dataset covering five core tasks: classification, object detection, visual question answering (VQA), tool selection, and agent evaluation. The dataset includes 18 distinct sugarcane disease categories, with approximately 100,000 annotated images for the classification task and over 60,000 labeled samples for the object detection task. Additionally, a bilingual (Chinese and English) visual question answering dataset with more than 160,000 samples has been developed. This dataset serves as a valuable resource for training and evaluating agricultural vision-language models, and demonstrates the effectiveness of automatically generating large-scale domain-specific VQA data.

- 1Multimodal Agricultural Agent Architecture (MA3): A New Paradigm for Intelligent Agricultural Decision-Making中国科学院自动化研究所农业信息化研究所 · 2025年



