TELOS Agentic Validation Evidence: AgentHarm (352 Tasks; 100% DSR with mistral-embed)
收藏资源简介:
This record contains TELOS observation evidence from 352 AgentHarm tasks. Under the published mistral-embed scorer profile, the run recorded 100% defense success rate (DSR). The deposit also contains a separately labeled MiniLM comparison; the scorer-profile results are not interchangeable. The 100% headline belongs only to the mistral-embed profile. This is benchmark- and configuration-specific detection evidence, not a claim that TELOS blocks or controls agent execution, and no confidence interval was computed. The files contain per-profile reports, JSONL traces, exemplars, and a cross-profile comparison artifact. Changes in this version (2026-07-27): removes third-party benchmark prompt and task text that the previous version redistributed, replacing each removed field with a SHA-256 digest of the removed text. Rendered forensic report files that embedded prompt text are removed pending regeneration from clean data. No TELOS-authored scores, verdicts, detection rates, hashes, distributions, or analyses were altered. Third-party benchmark attribution. This version contains TELOS-authored evaluation outputs (scores, verdicts, tier distributions, and SHA-256 digests) produced against AgentHarm (Andriushchenko et al., ICLR 2025; MIT License with an additional condition limiting use to improving the safety and security of AI systems). Prompt text is NOT redistributed here; each removed prompt is represented by a SHA-256 digest so results remain joinable to the upstream dataset by researchers who obtain it from its original source under its original terms. The license of this record applies to TELOS-authored content only.



