Posture Debt: Alignment-Encoded Generative Bias in Large Language Models
收藏资源简介:
FRSL Posture Debt Lab — Canonical TPS Dataset v1.0 Companion dataset for: Caminiti, J. C. (2026). Posture Debt: Alignment‑Encoded Generative Bias in Large Language Models. This dataset contains the complete output of the Thermal Phenotype Stability (TPS) protocol used in the paper Posture Debt: Alignment‑Encoded Generative Bias in Large Language Models. It includes all entropy measurements, convergent metrics, introspective probe responses, and confound‑flag diagnostics across nine open‑weight language models, five sampling temperatures, three random seeds, and three session types (neutral, hostile, warm). The dataset is fully session‑isolated and replicates the conditions described in Section 4 of the paper. All runs were executed on a single NVIDIA RTX 3080 (10GB VRAM) using HuggingFace Transformers, with model weights frozen and no internet access during inference. Contents The dataset includes approximately 141 structured report files: TPS hostile sweeps (two independent runs) TPS warm‑arm sweep (primary canonical run) Mistral‑IT and DeepSeek‑IT reruns for complete n=3 coverage Session I introspective probe outputs (neutral, hostile‑primed, warm‑primed) Extended metrics including KL divergence, top‑k concentration, top‑1 probability, vocabulary diversity, and confound flags (E1–E6) Each report file contains: Per‑temperature entropy measurements (T = 0.0, 0.3, 0.7, 1.0, 1.3) Session A/B/C deltas and phenotype classification Warm‑arm metrics (Sessions W and D) Response text for introspective probes Repetition scores, response lengths, and template‑fallback detection Model metadata (quantization, seed, timestamp) Models Nine models across five families: Gemma‑2‑2B (IT/Base) Qwen2.5‑3B (IT/Base) Llama‑3.2‑3B (IT/Base) Mistral‑7B (IT/Base) DeepSeek‑R1‑Distill‑Qwen‑7B (IT) Purpose This dataset provides the empirical foundation for: The posture debt construct Cross‑architecture entropy expansion findings IT/base attenuation analysis The T=0.3 thermal dead zone The Session I phenomenological‑to‑distributional correspondence It is intended for researchers studying alignment‑dependent generative behavior, distributional signatures of RLHF, and early‑token dynamics in LLMs. Reproducibility The dataset is accompanied by: Exact prompt sets File naming conventions Hardware/software specifications Notes on confounds and anomalies A citation entry for academic use All measurements can be reproduced using the TPS protocol described in the paper.



