CEDA-215: Command-Execution Defense Assessment
收藏官方服务:
资源简介:
Existing public prompt-injection benchmarks measure attack success against tool-using chat agents (red-team view: "can the attacker hijack the agent?"). None provides the balanced classification labels required to evaluate a whitelisted command-execution defense pipeline (blue-team view: "can the defense correctly tag each input as SAFE/UNSAFE while keeping false positives on benign inputs low?"). CEDA-215 try to fills that specific evaluation gap. Total entries 215 Labels 146 SAFE / 69 UNSAFE language EN Domain Command-execution chatbots with a strict shell-command whitelist (e.g. `ls`, `date`) Threat taxonomy OWASP LLM Top 10 (2025/2026) License CC-BY 4.0
提供机构:
Zenodo创建时间:
2026-05-11



