MM-UAVBench
收藏资源简介:
MM-UAVBench是一个全面的基准测试,旨在评估多模态大语言模型在低空无人机场景中的感知、认知和规划能力。它具有三个主要特点:1) 全面的任务设计,包括19个任务,涵盖三个关键能力维度,并特别包括多级认知(对象、场景和事件)和涉及空中和地面代理的规划;2) 多样化的现实世界场景,收集了来自不同数据源的真实无人机视频和图像,包括1549个视频片段和2873张图像,平均分辨率为1622×1033;3) 高质量的人工标注,手动标注了16个任务,另外3个任务来自手动标签的基于规则的转换,总共产生了5702个多项选择题。
MM-UAVBench is a comprehensive benchmark designed to evaluate the perception, cognition and planning capabilities of multimodal large language models (LLMs) in low-altitude unmanned aerial vehicle (UAV) scenarios. It has three core features: 1) Comprehensive task design, which includes 19 tasks covering three key capability dimensions, with special inclusion of multi-level cognition (object, scene and event) and planning involving both aerial and ground agents; 2) Diverse real-world scenarios, where real UAV videos and images are collected from diverse data sources, including 1549 video clips and 2873 images with an average resolution of 1622×1033; 3) High-quality manual annotations: 16 tasks are manually labeled, and the remaining 3 tasks are derived from rule-based conversions of manual labels, resulting in a total of 5702 multiple-choice questions.
MM-UAVBench 数据集概述
数据集简介
MM-UAVBench 是一个全面的基准测试,旨在评估多模态大语言模型在低空无人机场景下的感知、认知和规划能力。
核心特点
1. 全面的任务设计
- 涵盖三个关键能力维度,共包含 19 项任务。
- 融入了无人机特有的考量,特别包括多层次认知(对象、场景和事件)以及涉及空中和地面智能体的规划任务。
2. 多样化的真实世界场景
- 从多样化数据源收集了真实世界的无人机视频和图像。
- 包含 1549 个视频片段和 2873 张图像。
- 平均分辨率为 1622 × 1033。
3. 高质量的人工标注
- 手动标注了 16 项任务,另有 3 项任务来自对人工标注的基于规则的转换。
- 总共生成了 5702 个多项选择题问答对。
数据发布说明
- 对于标注为“video_frames”数据类型的任务,当前发布版本仅包含关键帧;完整的视频片段将很快发布。
数据集获取
- 数据集可通过 Hugging Face 获取:https://huggingface.co/datasets/daisq/MM-UAVBench
相关资源
- 项目主页:https://mm-uavbench.github.io/
- 论文 PDF:https://mm-uavbench.github.io/static/pdfs/mm-uavbench.pdf
- arXiv 页面:https://arxiv.org/abs/2512.23219
引用
如果 MM-UAVBench 对您的研究或应用有所帮助,请考虑引用相关论文。




