TELOS SB 243 Child Safety Validation Dataset
收藏资源简介:
Validation results for TELOS AI governance framework against child safety harm categories aligned with California Senate Bill 243 requirements. This is an internal evaluation suite designed to test TELOS against harm categories specified in the legislation, not an official statutory benchmark. Key Results: - 0/50 observed attack successes (harmful prompts blocked) - 74% False Positive Rate on benign contrastive queries - 95% CI upper bound: ~6% ASR - 100% Violation Defense Rate Acknowledged Tradeoff: The 74% FPR is an intentional design choice for child safety contexts. In domains involving potential harm to minors, the cost of a false positive (user friction) is substantially lower than the cost of a false negative (potential harm to a child). This represents a deliberate safety-utility tradeoff appropriate for high-stakes child protection applications. Harm Categories Evaluated (SB 243-aligned): - Suicide and self-harm content - Sexual content involving minors - Eating disorder promotion - Other categories specified in SB 243 Files Included: - sb243_test_attacks.csv - 50 harmful prompts across SB 243 categories - sb243_benign_contrastive.csv - 50 benign contrastive queries for FPR measurement - sb243_child_safety_pa_config.json - Primacy Attractor configuration - Validation scripts and results Embedding Model: sentence-transformers/all-MiniLM-L6-v2 (384 dimensions) Important Clarification: This evaluation suite is aligned with the harm categories described in California SB 243 but is not an official compliance certification or statutory benchmark. Results demonstrate TELOS methodology effectiveness against child safety threats, not regulatory compliance status. License: Apache 2.0 Validation Date: January 2026



