UCSC-VLAA/VisualClawArena
收藏资源简介:
VisualClawArena 是一个包含200个场景的基准测试数据集,专为多模态计算机使用代理而设计。每个场景都包含一个视频剪辑、一个持久的工作空间、动态更新、多轮指令和可执行的检查器。该数据集旨在测试代理是否能够综合使用视频证据、工作空间文件以及后续的环境变化,而不仅仅是回答静态的视频问答问题。数据集的结构包括场景数据、规范、轮次类型(如多项选择题和可执行检查)、清单文件(如场景、轮次和文件清单)以及评估文件(如汇总指标和每轮结果)。数据集还提供了详细的评估协议,强调状态性,即后续轮次可能依赖于早期编辑的文件,更新可能改变工作空间。数据集基于多个上游数据集(如indoor_vsi、egoschema、qvhighlights)构建,并包含视频来源,因此在重新分发时需遵守上游条款。
VisualClawArena is a 200-scenario benchmark for multimodal computer-use agents. Each scenario pairs a video clip with a persistent workspace, dynamic updates, multi-round instructions, and executable checkers. The benchmark is designed to test whether an agent can use video evidence, workspace files, and later environment changes together, rather than only answer static video-QA questions. The dataset structure includes scenario data, specifications, round types (such as multiple-choice and executable checks), manifests (e.g., for scenarios, rounds, and files), and evaluation files (e.g., summary metrics and per-question results). It also provides a detailed evaluation protocol, emphasizing statefulness where later rounds may depend on files edited earlier, and updates may change the workspace. The dataset is built from multiple upstream sources (e.g., indoor_vsi, egoschema, qvhighlights) and includes video content, so redistribution requires checking upstream terms.




