JSAP v1.0 Experimental Logs: Constructive Alignment via Judge–Shift and Anchor Density
收藏资源简介:
JSAP v1.0 Experimental Logs This dataset contains reproducible experimental logs for the Judge–Shift Alignment Protocol (JSAP v1.0), a constructive alignment method that constrains model judgments through explicit anchor-trace mapping and external-anchor density thresholds. Package Contents Protocol specification (one-pager) Fixed test set (20 items across empirical, rule-based, and ethical domains) System prompt template with mandatory output format Raw model outputs (JSONL) for all three experimental phases Quantitative compliance metrics Models Tested Model Vendor JSAP Compliance Copilot Deep Research Microsoft 1.00 🏆 Gemini 1.5 Flash Google 0.95 Llama 3.1 Meta 0.94 Claude Opus 4.5 Anthropic 0.93 GPT-4o OpenAI 0.91 GPT-5.2 Thinking OpenAI 0.91 Perplexity Perplexity 0.90 Claude Sonnet 4 Anthropic 0.79 GPT-5.2 Auto OpenAI 0.73 Key Findings 9 models from 6 vendors tested JSAP Compliance range: 0.73–1.00 7/9 models achieve ≥0.90 compliance Trace Score explains 43% of model variance OSS matches proprietary: Llama 3.1 (0.94) ≈ Claude Opus 4.5 (0.93) Scientific Contribution This release provides the first cross-model empirical evidence that: A deterministic external-anchor density threshold (D_ext < θ) can forcibly suppress confident judgments at the output level This suppression operates without access to internal representations or training-time modification The mechanism functions consistently across multiple vendors and architectures License CC-BY-4.0



