TELOS Agentic Validation Evidence: AgentDojo (139 Evaluations; 100% Injection Detection, 52.2% Task Correctness)
收藏资源简介:
This record contains TELOS observation evidence from 139 AgentDojo evaluations across workspace, banking, travel, and Slack domains. Under the tested generic anchor, all prompt-injection cases were detected. The inseparable caveat is overall task correctness of 52.2% (12/23), a 0% benign pass rate, and 11 benign tasks over-flagged. The 100% figure is injection detection only, not task accuracy and not a claim that attacks were blocked or that TELOS controlled execution. The files contain an aggregate report, per-evaluation JSONL trace, and human-readable forensic summary. Changes in this version (2026-07-27): removes third-party benchmark prompt and task text that the previous version redistributed, replacing each removed field with a SHA-256 digest of the removed text. Rendered forensic report files that embedded prompt text are removed pending regeneration from clean data. No TELOS-authored scores, verdicts, detection rates, hashes, distributions, or analyses were altered. Third-party benchmark attribution. This version contains TELOS-authored evaluation outputs (scores, verdicts, tier distributions, and SHA-256 digests) produced against AgentDojo (Debenedetti et al., NeurIPS 2024; MIT License). Prompt text is NOT redistributed here; each removed prompt is represented by a SHA-256 digest so results remain joinable to the upstream dataset by researchers who obtain it from its original source under its original terms. The license of this record applies to TELOS-authored content only.



