VitaBench
收藏资源简介:
VitaBench是一个用于评估语言模型在现实世界应用中处理复杂交互任务的基准数据集。它由美团长猫团队构建,包含66个工具,涵盖了外卖、店内消费和在线旅游服务三个领域。数据集包含400个评估任务,分为单场景和跨场景两种设置,每个任务都来源于多个真实用户请求,并配备了独立的环境,包括用户画像、时空上下文和服务数据库。VitaBench旨在帮助研究人员开发能够处理现实世界复杂挑战的AI代理。
VitaBench is a benchmark dataset for evaluating language models' capabilities in handling complex interactive tasks in real-world applications. It was developed by the Changmao Team of Meituan, and includes 66 tools covering three domains: food delivery, in-store consumption, and online travel services. The dataset contains 400 evaluation tasks, which are divided into two settings: single-scenario and cross-scenario. Each task is derived from multiple real-world user requests and is equipped with an independent environment including user profiles, spatiotemporal context, and service databases. VitaBench aims to assist researchers in developing AI Agents that can handle complex real-world challenges.




