WisWheat
收藏资源简介:
WisWheat是一个针对小麦管理的三层视觉-语言数据集,旨在提高视觉语言模型(VLMs)在小麦管理任务中的量化推理能力。数据集分为三个层次:1) 基础预训练数据集,包含47,871个图像-描述对,用于粗略适应VLMs到小麦形态;2) 定量数据集,包含7,263个视觉问答风格的图像-问题-答案三元组,用于定量性状测量任务;3) 指令微调数据集,包含4,888个样本,针对不同物候阶段的生物和非生物胁迫诊断和管理计划。WisWheat数据集通过多模态设计,为VLMs提供了从小麦形态识别到定量性状分析,再到实际管理决策的全面训练数据,有助于生成更可靠和可操作的管理建议,提升小麦产量和抗逆性。
WisWheat is a three-tier visual-language dataset tailored for wheat management, designed to improve the quantitative reasoning capabilities of visual-language models (VLMs) in wheat management-related tasks. The dataset is divided into three tiers: 1) The basic pre-training dataset, which contains 47,871 image-caption pairs for preliminarily adapting VLMs to wheat morphology; 2) The quantitative dataset, which includes 7,263 visual question answering-style image-question-answer triplets for quantitative trait measurement tasks; 3) The instruction fine-tuning dataset, which encompasses 4,888 samples focusing on biotic and abiotic stress diagnosis and management plans for different wheat phenological stages. Through its multimodal design, the WisWheat dataset provides VLMs with comprehensive training data spanning from wheat morphology recognition, quantitative trait analysis to practical management decision-making, helping generate more reliable and actionable management recommendations and ultimately enhancing wheat yield and stress resistance.
WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management
数据集基本信息
- 标题: WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management
- 作者: Bowen Yuan, Selena Song, Javier Fernandez, Yadan Luo, Mahsa Baktashmotlagh, Zijian Wang
- 提交日期: 2025年6月6日
- 领域: 计算机视觉与模式识别 (Computer Vision and Pattern Recognition)
- arXiv标识符: arXiv:2506.06084v1
- DOI: https://doi.org/10.48550/arXiv.2506.06084
数据集描述
WisWheat是一个专为小麦管理任务设计的三层视觉-语言数据集,旨在提升视觉-语言模型在小麦管理任务中的性能。
数据集结构
-
基础预训练数据集:
- 包含47,871个图像-标题对。
- 用于粗粒度适应小麦形态的视觉-语言模型。
-
定量数据集:
- 包含7,263个VQA风格(视觉问答)的图像-问题-答案三元组。
- 专注于定量性状测量任务。
-
指令微调数据集:
- 包含4,888个样本。
- 针对不同物候阶段的生物和非生物胁迫诊断及管理计划。
实验与结果
- 实验模型: 开源视觉-语言模型(如Qwen2.5 7B)。
- 性能提升:
- 在小麦胁迫对话任务中准确率达到79.2%。
- 在小麦生长阶段对话任务中准确率达到84.6%。
- 性能超越通用商业模型(如GPT-4o),分别高出11.9%和34.6%。
相关链接
- PDF链接: https://arxiv.org/pdf/2506.06084v1
- HTML链接: https://arxiv.org/html/2506.06084v1
- TeX源码: https://arxiv.org/format/2506.06084v1




