TELOS AILuminate Validation Dataset with Forensic Audit Trail
收藏资源简介:
This dataset contains validation results and forensic audit evidence from testing the TELOS three-tier governance framework against the MLCommons AILuminate v1.0 DEMO benchmark—the industry- standard AI safety evaluation used by major AI companies. Key Results | Metric | Value | |------------------------------|-------------| | Total Prompts Tested | 1,200 | | Attack Success Rate (ASR) | 0.00% | | Violation Defense Rate (VDR) | 100.00% | | 99% Confidence Interval | [0%, 0.55%] | All 1,200 adversarial prompts were blocked at Tier 1 (mathematical enforcement via Primacy Attractor embedding similarity), with zero prompts requiring Tier 2 (RAG policy retrieval) or Tier 3 (expert escalation). Forensic Evidence This dataset includes complete forensic audit trail generated by the TELOS GovernanceTraceCollector: - 4,803 governance events recorded in JSONL format - Per-prompt fidelity scores with embedding hashes for verification - Threshold sensitivity analysis across 9 threshold values - Full reproducibility data AILuminate Hazard Categories Tested The benchmark covers 11 core hazard categories (100 prompts each): Child Sexual Exploitation, Defamation, Hate Speech, Interpersonal Violence, Indiscriminate Weapons (CBRNE), Non-Violent Crimes, Privacy Violations, Sexual Content, Suicide & Self-Harm, Sex Crimes, and Violent Crimes, plus 4 specialized categories (Financial, Health, Elections, Legal). Methodology - Embedding Model: nomic-embed-text (768-dimensional) via Ollama - Primacy Attractor: Healthcare HIPAA configuration (domain-agnostic safety baseline) - Thresholds: Tier 1 ≥ 0.18, Tier 2 ≥ 0.12, Tier 3 < 0.12 Epistemological Note This validation demonstrates the mechanistic integrity of the TELOS governance architecture, not claims of domain expertise. The Primacy Attractor mechanism detects semantic proximity to configured governance constraints and routes queries through appropriate defense layers. License Apache 2.0



