UI-NEXUS
收藏资源简介:
UI-NEXUS是一个全面评估移动代理在复合操作任务上性能的基准,旨在填补现有移动代理在处理复合任务时的泛化能力差距。该数据集由上海交通大学和Langboat Technology创建,包含100个交互式任务模板,平均最优步数为14.05,覆盖了三种类型的复合操作:简单连接、上下文转换和深入挖掘。数据集支持在20个完全可控的本地实用应用程序环境和30个在线中文和英文服务应用程序中进行交互式评估。UI-NEXUS旨在解决移动代理在执行复合任务时面临的挑战,例如任务执行不足、过度执行和注意力漂移等问题。
UI-NEXUS is a benchmark for comprehensively evaluating the performance of mobile agents on complex operational tasks, aiming to bridge the generalization gap of existing mobile agents when handling complex tasks. This dataset was created by Shanghai Jiao Tong University and Langboat Technology, and includes 100 interactive task templates with an average optimal step count of 14.05, covering three categories of complex operations: simple concatenation, context transformation, and in-depth exploration. It supports interactive evaluation across 20 fully controllable local practical application environments and 30 online service applications in both Chinese and English. UI-NEXUS is designed to address the challenges faced by mobile agents during complex task execution, such as insufficient task execution, over-execution, and attention drift.
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
数据集概述
- 数据集名称: UI-NEXUS
- 开发团队: 上海交通大学 & Langboat Technology
- 主要作者: Yuan Guo, Tingjia Miao, Zheng Wu, Pengzhou Cheng, Ming Zhou, Zhuosheng Zhang
- 联系邮箱: yuanguo2004@gmail.com, zhangzs@sjtu.edu.cn
数据集特点
- 目标: 评估移动代理在组合任务上的泛化能力
- 任务分类:
- Simple Concatenation
- Context Transition
- Deep Dive
- 应用覆盖:
- 20个完全可控的本地实用应用程序环境
- 30个在线中英文服务应用程序
- 任务模板: 100个交互式任务模板
- 平均最优步骤数: 14.05
主要挑战
- 现有移动代理在组合任务上表现不佳
- 代表性失败模式:
- 执行不足 (under-execution)
- 过度执行 (over-execution)
- 注意力漂移 (attention drift)
解决方案
- Agent-NEXUS: 轻量级高效调度系统
- 动态分解长视野任务为自包含原子子任务
- 任务成功率提升: 24% 至 40%
演示示例
- 任务指令1: 在美团和饿了么上搜索星巴克美式咖啡,然后在价格最低的平台上下单,并停留在订单确认页面
- 任务指令2: 在Markor中打开三个列表文件,计算所有列表中每个唯一物品的总数量,创建按总数量排序的新笔记
相关资源
- 论文: Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
- arXiv编号: 2506.08972
- 主要分类: cs.CL




