SolshineMisfit/nla-qwen2.5-7b-L20-results-famous-ads
收藏资源简介:
该数据集名为著名广告标语——NLA激活分析(Qwen2.5-7B,第20层),旨在探究语言模型在读取著名广告标语时的内部表示。它使用自然语言自编码器(NLA)将Qwen2.5-7B-Instruct模型在第20层最后一个token的残差流激活解码为自然语言描述。数据集包含约146条著名广告标语(源自维基百科,附有品牌、年份和上下文信息),每条标语都对应一个由AV生成的英文解释(描述模型对该标语的表征内容)以及AR计算的余弦相似度和均方误差保真度分数。该数据集是一个针对广告语言的小型可解释性探针,用于研究模型如何编码压缩、高识别度的营销短语,适用于可解释性和机制可解释性研究。数据集基于CC BY-SA 4.0许可证发布,标语文本和上下文来自维基百科,nla_explanation由公开的NLA模型生成。
This dataset, titled Famous Ad Slogans — NLA Activation Analysis (Qwen2.5-7B, Layer 20), explores what a language model internally represents when reading famous advertising slogans. It uses a Natural Language Autoencoder (NLA) to decode the residual-stream activation at the final token of layer 20 in the Qwen2.5-7B-Instruct model into plain English descriptions. The dataset contains ~146 famous slogans (sourced from Wikipedia, with brand, year, and context), each paired with an English explanation generated by an AV (Activation Verbalizer) describing the models representation of the slogan, along with cosine similarity and MSE fidelity scores computed by an AR (Activation Reconstructor). It serves as a small interpretability probe over advertising language, useful for studying how the model encodes compressed, high-recognition marketing phrases, and is intended for interpretability and mechanistic interpretability research. The dataset is released under the CC BY-SA 4.0 license, with slogan text and context from Wikipedia, and nla_explanation generated from released NLA models.



