MobileAgentBench
收藏资源简介:
MobileAgentBench是由卡内基梅隆大学等机构创建的一个高效且用户友好的移动LLM代理基准测试数据集。该数据集包含100个任务,分布在10个开源应用中,旨在通过模拟日常任务来评估移动代理的性能。数据集的创建过程注重任务的成功条件灵活性和低代码侵入性,使得第三方开发者能够轻松扩展和定制任务。MobileAgentBench的应用领域广泛,主要用于学术和工业界,以解决移动代理性能评估的挑战,推动智能个人助理技术的发展。
MobileAgentBench is an efficient and user-friendly mobile LLM agent benchmark dataset created by Carnegie Mellon University and other institutions. It consists of 100 tasks distributed across 10 open-source applications, aiming to evaluate the performance of mobile agents by simulating real-world daily tasks. The dataset's creation process emphasizes flexible task success conditions and low code invasiveness, enabling third-party developers to easily extend and customize the tasks. With wide applicability, MobileAgentBench is primarily used in academic and industrial fields to address the challenges of mobile agent performance evaluation and promote the development of intelligent personal assistant technologies.




