osunlp/SkillHarm
收藏资源简介:
代理技能在代理工作流程中占据特权地位——代理被期望隐式遵循并执行它们——这使得第三方技能成为脆弱的供应链攻击面。SkillHarm是一个基于技能的跨技能使用生命周期攻击的基准,并配有一个包含12种技能相关风险的系统分类法。每个攻击都基于可运行的Harbor任务环境,并通过确定性的攻击成功率评估器进行评分。该基准包含879个攻击样本,涵盖两种攻击场景:固定载荷投毒(FPP)和自我突变投毒(SMP)。
Agent skills occupy a privileged position in the agent workflow — agents are expected to implicitly follow and execute them — which makes third-party skills a vulnerable supply-chain attack surface. SkillHarm is a benchmark of skill-based attacks across the skill-use lifecycle, paired with a systematic taxonomy of 12 skill-relevant risks. Every attack is grounded in a runnable Harbor task environment and scored by a deterministic attack-success-rate (ASR) evaluator. The benchmark contains 879 attack samples across two attack scenarios: Fixed-Payload Poisoning (FPP) and Self-Mutating Poisoning (SMP).




