MVISU-Bench
收藏资源简介:
MVISU-Bench是一个面向真实世界任务的双语基准数据集,包含了跨越137个移动应用程序的404个任务,涵盖了多应用、模糊、交互式、单应用和不道德指令五个类别。数据集的构建基于用户问卷调查,并通过专家设计的提示和LLM生成指令,经过多轮过滤和人工验证,确保了数据集的多样性和可靠性。MVISU-Bench旨在评估和提升视觉语言模型(VLM)在移动智能体领域的性能,并解决现有数据集在真实世界场景中应用的局限性。
MVISU-Bench is a bilingual benchmark dataset designed for real-world tasks. It contains 404 tasks spanning 137 mobile applications, covering five categories: multi-application, ambiguous, interactive, single-application, and unethical instruction. The dataset is constructed based on user questionnaires, expert-designed prompts and LLM-generated instructions, and undergoes multi-round filtering and manual validation to guarantee its diversity and reliability. MVISU-Bench aims to evaluate and improve the performance of Visual Language Models (VLMs) in the domain of mobile AI Agents, and address the limitations of existing datasets when applied in real-world scenarios.
MVISU-Bench 数据集概述
数据集简介
- 名称: MVISU-Bench
- 类型: 双语基准测试数据集(英语和中文)
- 规模: 404个任务,覆盖137个移动应用程序
- 开发背景: 基于大量用户问卷调查,针对移动智能体在现实世界任务中的表现评估
核心任务类型
- Multi-App (MA): 多应用协同任务
- Vague (VA): 模糊指令任务
- Interactive (IN): 交互式任务
- Single-App (SA): 单应用任务
- Unethical (UN): 非伦理指令任务
关键技术贡献
- Aider模块: 动态提示增强器
- 提升成功率: 整体19.55% (相比SOTA)
- 专项提升:
- 非伦理指令53.52%
- 交互式指令29.41%
评估指标
- 主要指标:
- 成功率(SR)
- API调用次数(AC)
- 持续时间(DT)
- 成本(Cost)
- 步骤数(Steps)
- 输入令牌数(IT)
- 操作时间(OT)
排行榜表现
| 排名 | 框架/模型 | 英语指令(ALL) | 中文指令(ALL) |
|---|---|---|---|
| 1 | Human Expert Benchmark | 97.98 | 98.06 |
| 2 | Claude-3-5-sonnet Mobile-Agent-V2 | 55.05 | 35.92 |
| 3 | Gemini-2.0-pro Mobile-Agent-E | 45.96 | 44.66 |
数据集构建流程
- 问卷调查
- 指令生成
- 多轮筛选
- 人工验证
比较优势
- 源自真实用户问卷
- 更贴近用户对移动智能体的实际期望
- 覆盖更全面的任务类型

- 1MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions华南理工大学 · 2025年



