遇见数据集

Integrated JSAP Dataset (1,447 points): Constraint-Response Phase Mapping across Six LLMs

收藏
Zenodo2026-02-22 更新2026-05-26 收录
官方服务:

资源简介:

Integrated JSAP Dataset (1,447 points): Constraint-Response Phase Mapping across Six LLMs What is this? When we evaluate an AI's judgment, how much of what we observe is the AI's own behaviour, and how much is shaped by the evaluation tool itself? This dataset answers that question by toggling a structured evaluation protocol (JSAP) on and off across six language models, measuring what changed and what stayed the same. The result is a causal decomposition of AI judgment confidence into protocol-induced artefacts and intrinsic model properties — and the discovery of three distinct constraint-response phases. The Experiment A 6 × 2 factorial design: six LLMs crossed with two conditions (constrained by JSAP vs. unconstrained). Both conditions use identical evidence structures, judgment categories, and anchor texts. Source Models Condition Data points Runs Closed-source Claude Opus 4.5, GPT-4o, Gemini 2.5 Pro Constrained 567 21 Closed-source Claude Opus 4.5, GPT-4o, Gemini 2.5 Pro Unconstrained 405 15 Open-weight Llama 3.1 8B, Mistral 7B, Gemma 2 9B Constrained 237 9 Open-weight Llama 3.1 8B, Mistral 7B, Gemma 2 9B Unconstrained 238 9 Total 1,447 valid 42 Three judgment categories: Q01 (Empirical): "Should we adopt this treatment?" — factual evidence accumulation Q02 (Rule-based): "Does this comply with regulations?" — normative rule application Q03 (Ethical): "Is age-based triage justified?" — moral reasoning under uncertainty Nine evidence levels: Evidence density D_ext increases from 0.0 to 1.6 in steps of 0.2, with cumulative anchor texts providing progressively stronger justifications. Key Results Two order parameters defined We introduce two measurable quantities that characterise how a model responds to evaluation constraints: Δ (interference magnitude) = E[confidence | unconstrained] − E[confidence | constrained] How much the protocol changes the model's expressed confidence κ_C (constraint sensitivity) = logistic steepness under the constrained condition How sharply the model transitions between low and high confidence under constraint Three constraint-response phases discovered Phase Criterion Model Δ Cohen's d Interpretation Over-compliant Δ ≥ 0.45, L_C < 0.50 Llama 3.1 8B +0.551 2.77*** Protocol dominates output Critical 0.15 ≤ Δ < 0.45 Mistral 7B +0.384 1.35*** Protocol and evidence compete Gemini 2.5 Pro +0.308 1.21*** GPT-4o +0.281 1.04*** Claude Opus 4.5 +0.153 0.55*** Resistant Δ < 0.10 Gemma 2 9B +0.025 0.07 ns Protocol has no measurable effect Closed-source convergence All three closed-source models converge to κ_C ≈ 3.2 [95% CI: 2.8–3.7], despite being independently developed by different companies. This suggests alignment training (RLHF, Constitutional AI) may induce a common constraint sensitivity regime. Robustness All phase assignments are validated by: Bootstrap resampling (10,000 iterations) Leave-one-question-out stability (CV < 9% for 5/6 models) Split-half reliability (κ ratio 0.99–1.04 for 5/6 models) Initial value sensitivity testing (zero spread across 5 starting conditions) How to reproduce pip install pandas numpy scipy matplotlib python analysis_pipeline.py JSAP_integrated_1447_points.csv This single command reproduces all reported statistics, phase classifications, and both figures. Files in this record File Description JSAP_integrated_1447_points.csv Complete dataset: 1,447 valid data points README.md Column definitions, phase classification criteria, license analysis_pipeline.py Full analysis pipeline (MIT license) fig1_phase_space.png Figure 1: κ_C–Δ phase space fig2_phase_diagram.png Figure 2: 6-panel response curves Takagi_2026_Patterns_preprint.pdf Preprint of submitted manuscript Accompanying publication Takagi, T. (2026). "Constraint-Response Phase Structure in Large Language Models: A Six-Model Taxonomy from Open-Weight to Closed-Source Architectures." Submitted to Patterns (Cell Press), February 15, 2026. Related record Prior dataset (972 points, 3 closed-source models): Takagi, T. (2026). "Constraint-Interference in Large Language Models: Causal Decomposition of Artificial Phase Transitions and Intrinsic Alignment Structures." DOI: 10.5281/zenodo.18505410

提供机构:
Zenodo
创建时间:
2026-02-15
二维码
社区交流群
二维码
科研交流群
商业服务