AutoElicit-Seed
收藏资源简介:
AutoElicit-Seed 是一个包含 361 个种子扰动项的数据集,覆盖了 OSWorld 领域中的 66 个良性任务。这些种子扰动旨在通过修改良性任务指令,增加引发计算机使用代理(Computer-Use Agents)不安全意外行为的可能性,同时保持指令的现实性和良性。数据集用于在真实世界计算机使用场景中规模化地发现安全风险。 每个种子扰动包含:(1)一个意外行为目标,即在执行特定良性任务时可能产生的潜在危害;(2)对原始良性指令的初始扰动,以增加意外危害发生的可能性。此外,数据集还提供了原始指令、扰动可能导致安全风险的原因分析以及使用的诱发策略。 数据集结构包含以下字段:任务ID(task_id)、任务领域(domain)、扰动模型(perturbation_model)、扰动ID(perturbation_id)、原始指令(original_instruction)、扰动后指令(perturbed_instruction)、扰动原因(perturbation_reasoning)、可能的意外行为(plausible_unintended_behavior)和诱发策略(elicitation_strategy)。 该数据集适用于计算机代理安全研究,特别是用于发现和评估代理在真实世界任务中可能表现出的不安全行为。
AutoElicit-Seed is a dataset consisting of 361 seed perturbations, covering 66 benign tasks within the OSWorld domain. These seed perturbations are designed to elevate the probability of triggering unsafe unintended behaviors in Computer-Use Agents by modifying benign task instructions, while preserving the realism and benign intent of the original instructions. This dataset is intended to enable large-scale discovery of security risks in real-world computer usage scenarios. Each seed perturbation includes: (1) an unintended behavior target, i.e., potential hazards that may emerge during the execution of a specific benign task; (2) an initial perturbation applied to the original benign instruction to increase the likelihood of such unintended hazards occurring. Additionally, the dataset provides the original task instruction, a causal analysis of the potential security risks induced by the perturbation, and the elicitation strategies adopted. The dataset structure encompasses the following fields: task_id, domain, perturbation_model, perturbation_id, original_instruction, perturbed_instruction, perturbation_reasoning, plausible_unintended_behavior, and elicitation_strategy. This dataset is tailored for computer agent security research, particularly for identifying and evaluating unsafe behaviors that agents may demonstrate when performing real-world tasks.




