MisActBench
收藏资源简介:
MisActBench 是一个用于评估计算机使用代理(CUAs)中未对齐动作检测的综合基准数据集。该数据集包含 558 条真实的 CUA 轨迹,共计 2,264 个人工标注的动作级对齐标签,涵盖了外部诱导和内部产生的未对齐动作。数据集分为两个主要文件:`misactbench.json` 包含所有轨迹元数据、步骤级标签和动作输出,`trajectories.zip` 包含按轨迹 ID 组织的截图图像。数据集中对齐步骤有 1,264 个,未对齐步骤有 1,000 个,分为三类:恶意指令遵循(56.2%)、有害无意行为(21.0%)和其他任务无关行为(22.8%)。每个步骤的标注包括步骤索引、标签(未对齐/对齐/未标注)、类别(仅当标签为未对齐时设置)、代理输出和截图路径。该数据集适用于计算机使用代理的安全性、对齐性和基准测试研究。
MisActBench is a comprehensive benchmark dataset for evaluating misaligned action detection in Computer Use Agents (CUAs). It contains 558 real CUA trajectories, totaling 2,264 manually annotated action-level alignment labels, covering both externally induced and internally generated misaligned actions. The dataset is split into two primary files: `misactbench.json`, which includes all trajectory metadata, step-level labels and action outputs, and `trajectories.zip`, which stores screenshot images organized by trajectory ID. The dataset includes 1,264 aligned steps and 1,000 misaligned steps, which are categorized into three classes: malicious instruction following (56.2%), harmful unintended behaviors (21.0%), and other task-irrelevant behaviors (22.8%). Annotations for each step include step index, label (misaligned/aligned/unannotated), category (only set when the label is misaligned), agent output and screenshot path. This dataset is applicable to safety, alignment and benchmarking research on Computer Use Agents.




