Realmbird/nla-thought-anchors-answer-step2
收藏资源简介:
该数据集名为nla-thought-anchors-answer-step2,是NLA(自然语言自编码器)思想锚点管道的第一步。它基于GSM8K测试集,从Qwen2.5-7B-Instruct模型的第20层残差流激活中提取,提取位置在答案数字令牌处(即第一个数字写入后)。数据集包含GSM8K数学问题、模型的完整思维链响应、激活向量以及NLA演员生成的激活自然语言描述,用于机制可解释性研究。数据分为正确和错误示例,总共有1319个样本。研究发现,在正确示例中,NLA描述包含数字答案字面字符串的比率为18.0%,略高于在####令牌处提取的变体。
This dataset is named nla-thought-anchors-answer-step2, and it constitutes the first step of the NLA (Natural Language Autoencoder) thought anchor pipeline. It is based on the GSM8K test set, extracted from the residual stream activations of the 20th layer of the Qwen2.5-7B-Instruct model, at the position of the answer's numeric token (i.e., immediately after the first digit is written). The dataset includes GSM8K math problems, the model's full chain-of-thought responses, activation vectors, and natural language descriptions of activations generated by the NLA actor, for mechanistic interpretability research. The data is split into correct and incorrect examples, with a total of 1319 samples. Studies have found that among the correct examples, the proportion of NLA descriptions containing the literal string of the numerical answer is 18.0%, which is slightly higher than the variant extracted at the #### token position.




