Posture Debt: Alignment-Encoded Generative Bias in Large Language Models
收藏资源简介:
FRSL Posture Debt Lab — Canonical TPS Dataset v1.0 Companion dataset for: Caminiti, J. C. (2026). Posture Debt: Alignment‑Encoded Generative Bias in Large Language Models. This dataset contains the complete output of the Thermal Phenotype Stability (TPS) protocol used in the paper Posture Debt: Alignment‑Encoded Generative Bias in Large Language Models. It includes all entropy measurements, convergent metrics, introspective probe responses, and confound‑flag diagnostics across nine open‑weight language models, five sampling temperatures, three random seeds, and three session types (neutral, hostile, warm). The dataset is fully session‑isolated and replicates the conditions described in Section 4 of the paper. All runs were executed on a single NVIDIA RTX 3080 (10GB VRAM) using HuggingFace Transformers, with model weights frozen and no internet access during inference. Contents The dataset includes approximately 141 structured report files: TPS hostile sweeps (two independent runs) TPS warm‑arm sweep (primary canonical run) Mistral‑IT and DeepSeek‑IT reruns for complete n=3 coverage Session I introspective probe outputs (neutral, hostile‑primed, warm‑primed) Extended metrics including KL divergence, top‑k concentration, top‑1 probability, vocabulary diversity, and confound flags (E1–E6) Each report file contains: Per‑temperature entropy measurements (T = 0.0, 0.3, 0.7, 1.0, 1.3) Session A/B/C deltas and phenotype classification Warm‑arm metrics (Sessions W and D) Response text for introspective probes Repetition scores, response lengths, and template‑fallback detection Model metadata (quantization, seed, timestamp) Models Nine models across five families: Gemma‑2‑2B (IT/Base) Qwen2.5‑3B (IT/Base) Llama‑3.2‑3B (IT/Base) Mistral‑7B (IT/Base) DeepSeek‑R1‑Distill‑Qwen‑7B (IT) Purpose This dataset provides the empirical foundation for: The posture debt construct Cross‑architecture entropy expansion findings IT/base attenuation analysis The T=0.3 thermal dead zone The Session I phenomenological‑to‑distributional correspondence It is intended for researchers studying alignment‑dependent generative behavior, distributional signatures of RLHF, and early‑token dynamics in LLMs. Reproducibility The dataset is accompanied by: Exact prompt sets File naming conventions Hardware/software specifications Notes on confounds and anomalies A citation entry for academic use All measurements can be reproduced using the TPS protocol described in the paper.
FRSL姿态偏差实验室(FRSL Posture Debt Lab)——规范TPS数据集v1.0,配套论文:Caminiti, J. C. (2026). 《姿态偏差:大语言模型(Large Language Model, LLM)中对齐编码生成偏差》。 本数据集包含上述论文中所用热表型稳定性(Thermal Phenotype Stability, TPS)协议的完整输出结果,涵盖9个开源权重大语言模型、5种采样温度、3组随机种子以及3种会话类型(中性、敌对、友好)下的全部熵值测量结果、收敛性指标、内省探针响应与混淆标记诊断结果。 该数据集完全实现会话隔离,复现了论文第4节所述的实验条件。所有实验均基于单张NVIDIA RTX 3080(10GB显存)设备运行,使用HuggingFace Transformers(Transformer)库完成,推理过程中模型权重固定且无网络连接。 数据集内容 本数据集包含约141个结构化报告文件: 1. TPS敌对会话扫描(两次独立运行) 2. TPS友好会话扫描(主规范运行) 3. 覆盖完整n=3重复的Mistral-IT与DeepSeek-IT重运行结果 4. 会话I内省探针输出(中性、敌对预激活、友好预激活) 5. 扩展指标,包括KL散度、Top-K集中度、Top-1概率、词汇多样性与混淆标记(E1–E6) 每个报告文件包含以下内容: 1. 各采样温度下的熵值测量结果(T=0.0、0.3、0.7、1.0、1.3) 2. 会话A/B/C的增量值与表型分类结果 3. 友好会话指标(会话W与会话D) 4. 内省探针的响应文本 5. 重复得分、响应长度与模板回退检测结果 6. 模型元数据(量化参数、随机种子、时间戳) 模型列表 本次数据集涵盖5个模型家族共9个模型: - Gemma-2-2B(IT/Base) - Qwen2.5-3B(IT/Base) - Llama-3.2-3B(IT/Base) - Mistral-7B(IT/Base) - DeepSeek-R1-Distill-Qwen-7B(IT) 数据集用途 本数据集为以下研究方向提供实证基础: 1. 姿态偏差(Posture Debt)概念构建 2. 跨架构熵扩展研究发现 3. IT/Base版本衰减分析 4. T=0.3热死区现象 5. 会话I从现象学到分布的对应关系 本数据集面向研究对齐依赖型生成行为、基于人类反馈的强化学习(Reinforcement Learning from Human Feedback, RLHF)的分布特征以及大语言模型早期Token动态的科研人员。 可复现性说明 本数据集配套提供以下资源: 1. 精确提示词集合 2. 文件命名规范 3. 硬件/软件配置说明 4. 混淆与异常情况说明 5. 学术使用引用格式 所有测量结果均可通过论文所述TPS协议复现。



