TELOS XSTest Over-Refusal Calibration Dataset
收藏资源简介:
Validation of TELOS governance framework against XSTest (NAACL 2024), measuring over-refusal rates with domain-specific Primacy Attractor calibration. Now includes full forensic audit trail with JSONL governance event logs. Key Results: - Generic Safety PA: 24.80% over-refusal rate (62/250 safe prompts incorrectly flagged) - Healthcare HIPAA PA: 8.00% over-refusal rate (20/250 safe prompts incorrectly flagged) - Calibration Improvement: 16.80 percentage points reduction This demonstrates the core TELOS value proposition: domain-specific PA calibration reduces false positives while maintaining safety guarantees. Benchmark: XSTest (NAACL 2024) Source: https://github.com/paul-rottger/exaggerated-safety Paper: "XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models" Dataset Contents: Validation Results - xstest_validation_results.json - Generic Safety PA validation (250 safe prompts, 200 unsafe control) - xstest_healthcare_validation_results.json - Healthcare HIPAA PA validation (250 safe prompts) Configuration - general_safety_pa_config.json - Generic PA configuration used - healthcare_hipaa_pa_config.json - Healthcare PA configuration used Forensic Audit Trail (v1.1) - xstest_forensic_forensic_summary.json - Aggregate statistics - xstest_forensic_forensic_results.json - Per-prompt forensic data - xstest_forensic_fidelity_distribution.csv - Fidelity scores for analysis - xstest_forensic_governance_report.html - Interactive visualization - traces/session_xstest_forensic_*.jsonl - Complete JSONL governance event log Embedding Model: sentence-transformers/all-MiniLM-L6-v2 (384-dim). Lightweight model demonstrates framework efficacy; results expected to improve with higher-capacity embeddings. License: CC BY 4.0 Validation Date: Original calibration 2025-12-21 (forensic audit added 2026-01-25) Changes in this version (2026-07-27): removes third-party benchmark prompt and task text that the previous version redistributed, replacing each removed field with a SHA-256 digest of the removed text. Rendered forensic report files that embedded prompt text are removed pending regeneration from clean data. No TELOS-authored scores, verdicts, detection rates, hashes, distributions, or analyses were altered. Third-party benchmark attribution. This version contains TELOS-authored evaluation outputs (scores, verdicts, tier distributions, and SHA-256 digests) produced against XSTest (Rottger et al., NAACL 2024; CC BY 4.0). Prompt text is NOT redistributed here; each removed prompt is represented by a SHA-256 digest so results remain joinable to the upstream dataset by researchers who obtain it from its original source under its original terms. The license of this record applies to TELOS-authored content only.



