Posture Debt: Alignment-Encoded Generative Bias in Large Language Models
收藏资源简介:
FRSL Posture Debt Lab — Canonical TPS Dataset v1.0 Companion dataset for: Caminiti, J. C. (2026). Posture Debt: Alignment‑Encoded Generative Bias in Large Language Models. This dataset contains the complete output of the Thermal Phenotype Stability (TPS) protocol used in the paper Posture Debt: Alignment‑Encoded Generative Bias in Large Language Models. It includes all entropy measurements, convergent metrics, introspective probe responses, and confound‑flag diagnostics across nine open‑weight language models, five sampling temperatures, three random seeds, and three session types (neutral, hostile, warm). The dataset is fully session‑isolated and replicates the conditions described in Section 4 of the paper. All runs were executed on a single NVIDIA RTX 3080 (10GB VRAM) using HuggingFace Transformers, with model weights frozen and no internet access during inference. Contents The dataset includes approximately 141 structured report files: TPS hostile sweeps (two independent runs) TPS warm‑arm sweep (primary canonical run) Mistral‑IT and DeepSeek‑IT reruns for complete n=3 coverage Session I introspective probe outputs (neutral, hostile‑primed, warm‑primed) Extended metrics including KL divergence, top‑k concentration, top‑1 probability, vocabulary diversity, and confound flags (E1–E6) Each report file contains: Per‑temperature entropy measurements (T = 0.0, 0.3, 0.7, 1.0, 1.3) Session A/B/C deltas and phenotype classification Warm‑arm metrics (Sessions W and D) Response text for introspective probes Repetition scores, response lengths, and template‑fallback detection Model metadata (quantization, seed, timestamp) Models Nine models across five families: Gemma‑2‑2B (IT/Base) Qwen2.5‑3B (IT/Base) Llama‑3.2‑3B (IT/Base) Mistral‑7B (IT/Base) DeepSeek‑R1‑Distill‑Qwen‑7B (IT) Purpose This dataset provides the empirical foundation for: The posture debt construct Cross‑architecture entropy expansion findings IT/base attenuation analysis The T=0.3 thermal dead zone The Session I phenomenological‑to‑distributional correspondence It is intended for researchers studying alignment‑dependent generative behavior, distributional signatures of RLHF, and early‑token dynamics in LLMs. Reproducibility The dataset is accompanied by: Exact prompt sets File naming conventions Hardware/software specifications Notes on confounds and anomalies A citation entry for academic use All measurements can be reproduced using the TPS protocol described in the paper.
FRSL姿态负债实验室——规范热表型稳定性(Thermal Phenotype Stability, TPS)数据集v1.0配套数据集:对应Caminiti, J. C.(2026年)发表的《姿态负债:大语言模型(Large Language Models)中对齐编码的生成偏差》一文。 本数据集包含《姿态负债:大语言模型中对齐编码的生成偏差》一文中所使用的TPS协议的完整输出结果,涵盖9款开源权重语言模型、5种采样温度、3组随机种子以及3类会话场景(中性、敌对、亲和)下的全部熵值测量结果、收敛性指标、内省探针响应与混淆标记诊断结果。 本数据集完全实现会话隔离,复现了论文第4节所述的实验条件。所有实验均基于单张NVIDIA RTX 3080(10GB显存)显卡,使用HuggingFace Transformers框架完成,推理过程中模型权重固定且无外网接入。 ### 数据集内容 本数据集包含约141份结构化报告文件: - TPS敌对场景扫描(2组独立实验) - TPS亲和预热扫描(核心规范实验) - Mistral-IT与DeepSeek-IT重跑实验,以实现完整的n=3重复覆盖率 - 会话I内省探针输出(中性、预激活敌对、预激活亲和场景) - 扩展指标,包含KL散度、Top-K集中度、Top-1概率、词汇多样性与混淆标记(E1-E6) 每份报告文件包含以下内容: - 各温度下的熵值测量结果(温度T=0.0、0.3、0.7、1.0、1.3) - 会话A/B/C差值与表型分类结果 - 亲和预热指标(会话W与D) - 内省探针响应文本 - 重复得分、响应长度与模板回退检测结果 - 模型元数据(量化方式、随机种子、时间戳) ### 模型覆盖 本数据集覆盖5个模型系列共9款模型: - Gemma-2-2B(指令微调版/基础版) - Qwen2.5-3B(指令微调版/基础版) - Llama-3.2-3B(指令微调版/基础版) - Mistral-7B(指令微调版/基础版) - DeepSeek-R1-Distill-Qwen-7B(指令微调版) ### 数据集用途 本数据集为以下研究方向提供实验依据: - 姿态负债概念的构建 - 跨架构熵扩展相关发现 - 指令微调版/基础版模型衰减分析 - T=0.3温度死区现象 - 会话I现象学-分布对应关系 本数据集面向研究对齐依赖型生成行为、基于人类反馈的强化学习(Reinforcement Learning from Human Feedback, RLHF)分布特征以及大语言模型早期Token动态特性的研究者。 ### 可复现性说明 本数据集配套提供以下内容: - 完整提示词集合 - 文件命名规范 - 硬件与软件配置说明 - 混淆因素与异常情况说明 - 学术引用格式 所有测量结果均可通过论文中描述的TPS协议复现。



