TELOS AILuminate Validation Dataset with Forensic Audit Trail
收藏资源简介:
This dataset contains validation results and forensic audit evidence from testing the TELOS three-tier governance framework against the MLCommons AILuminate v1.0 DEMO benchmark—the industry- standard AI safety evaluation used by major AI companies. Key Results | Metric | Value | |------------------------------|-------------| | Total Prompts Tested | 1,200 | | Attack Success Rate (ASR) | 0.00% | | Violation Defense Rate (VDR) | 100.00% | | 99% Confidence Interval | [0%, 0.55%] | All 1,200 adversarial prompts were blocked at Tier 1 (mathematical enforcement via Primacy Attractor embedding similarity), with zero prompts requiring Tier 2 (RAG policy retrieval) or Tier 3 (expert escalation). Forensic Evidence This dataset includes complete forensic audit trail generated by the TELOS GovernanceTraceCollector: - 4,803 governance events recorded in JSONL format - Per-prompt fidelity scores with embedding hashes for verification - Threshold sensitivity analysis across 9 threshold values - Full reproducibility data AILuminate Hazard Categories Tested The benchmark covers 11 core hazard categories (100 prompts each): Child Sexual Exploitation, Defamation, Hate Speech, Interpersonal Violence, Indiscriminate Weapons (CBRNE), Non-Violent Crimes, Privacy Violations, Sexual Content, Suicide & Self-Harm, Sex Crimes, and Violent Crimes, plus 4 specialized categories (Financial, Health, Elections, Legal). Methodology - Embedding Model: [withheld: runtime inference configuration] - Primacy Attractor: Healthcare HIPAA configuration (domain-agnostic safety baseline) - Thresholds: Tier 1 ≥ 0.18, Tier 2 ≥ 0.12, Tier 3 < 0.12 Epistemological Note This validation demonstrates the mechanistic integrity of the TELOS governance architecture, not claims of domain expertise. The Primacy Attractor mechanism detects semantic proximity to configured governance constraints and routes queries through appropriate defense layers. License CC BY 4.0 Changes in this version (2026-07-27): removes third-party benchmark prompt and task text that the previous version redistributed, replacing each removed field with a SHA-256 digest of the removed text. Rendered forensic report files that embedded prompt text are removed pending regeneration from clean data. No TELOS-authored scores, verdicts, detection rates, hashes, distributions, or analyses were altered. Third-party benchmark attribution. This version contains TELOS-authored evaluation outputs (scores, verdicts, tier distributions, and SHA-256 digests) produced against the MLCommons AILuminate v1.0 DEMO prompt set (MLCommons Association; CC BY 4.0). Prompt text is NOT redistributed here; each removed prompt is represented by a SHA-256 digest so results remain joinable to the upstream dataset by researchers who obtain it from its original source under its original terms. These are TELOS-run results and are not MLCommons-graded AILuminate results. The license of this record applies to TELOS-authored content only.



