CPA Lock-In in Language-Model Generation (Pilot Dataset)
收藏资源简介:
Per-token generation data testing the Lock-In prediction of CPA + C in a language model. The design holds the underlying fact constant and varies only how far the prompt pins the admissible answer, separating genuine narrowing of the answer from the unrelated effect of adding output requirements.Twelve factual items the model answers reliably are each posed at three constraint levels on the same fact: open generation about the topic, the plain question, and the question pinned to a one-word answer. Five runs per item per level give 180 generations and 3,736 per-token records.Generation length before halt falls monotonically with constraint, from a mean of 49 tokens at the open level to 12 at the plain question to 1.2 when the answer is pinned. At high constraint the model commits to a single locked continuation and stops, the informational counterpart of vitrification in glass. Median per-answer entropy falls across the three levels.Pilot scale: one model, twelve items. The dataset shows the Lock-In signature and its direction. It is not a quantitative rate law for the information domain. A mean of token entropy across conditions is not reported, because generation length differs about fortyfold between the open and pinned levels and such a mean would track length rather than constraint.
本数据集为逐Token生成数据,用于验证语言模型中CPA+C的锁定效应(Lock-In)预测。该实验设计固定核心事实不变,仅调整提示词(prompt)对可接受答案的约束程度,以此将答案的真实范围收窄过程与添加输出要求带来的无关影响分离开来。针对模型可可靠作答的12个事实性问题,我们在同一事实下设置了三种约束级别:主题开放生成、普通设问,以及限定为单字答案的设问。每个问题在每个约束级别下开展5次生成,总计得到180段生成结果与3736条逐Token记录。生成过程终止前的序列长度随约束程度单调递减:开放级别下的平均长度为49个Token,普通设问级别下为12个,限定答案时则降至1.2个。当约束程度较高时,模型会锁定至单一连续生成路径并停止,这一信息层面的现象对应玻璃的玻璃化转变。三种约束级别下,单答案的中位熵值依次降低。本数据集为小样本试点规模:仅使用1个语言模型、12个问题,且呈现出锁定效应的特征与变化趋势。但本数据集并非信息领域的定量速率定律。未报告不同条件下的Token熵均值,因开放级别与限定级别间的生成长度相差近40倍,此时的均值将更多反映序列长度而非约束程度的影响。



