CUARewardBench
收藏资源简介:
CUARewardBench 是一个用于评估计算机使用代理(CUA)奖励模型的新基准,包含来自 10 个软件类别和 7 种代理架构的轨迹,具有不同的性能水平。所有轨迹都经过专家精心设计的协议进行标注,并通过严格的质量控制确保可靠性和实际适用性。该数据集旨在解决现有奖励模型在视觉推理能力、知识不足和通用 VLM 与专用 CUA 模型之间的优劣比较等问题。
CUARewardBench is a novel benchmark for evaluating Computer Use Agent (CUA) reward models. It encompasses trajectories across 10 software categories and 7 agent architectures, spanning diverse performance levels. All trajectories are annotated following expert-curated protocols, and their reliability and practical applicability are ensured through strict quality control measures. This benchmark aims to resolve critical issues plaguing current reward models, including inadequate visual reasoning capabilities, insufficient domain knowledge, and the lack of rigorous comparative analysis between general-purpose VLMs and specialized CUA models.
- 1CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent腾讯优图实验室 · 2025年



