Anisotropy and Cross-Model Comparison of Released Natural Language Autoencoders
收藏资源简介:
Cross-model calibration data for Anthropic's released Natural Language Autoencoders (NLAs). This dataset accompanies an empirical study of the released NLA checkpoints from Fraser-Taliente et al. (2026) on two open base models: Qwen2.5-7B-Instruct (extraction layer 20) and Gemma-3-12B-IT (extraction layer 32). It contains the raw activations, NLA explanations, reconstruction metrics, hallucination judgments, and chain-of-thought (CoT) probe results from a 500-input cross-tier calibration experiment. Phase 1: Tiered reconstruction calibration (N=500 per model). Inputs are stratified across five tiers: factual English (A), code (B), Turkish (C), legal/formal (D), and adversarial, with the adversarial tier split into structural anomaly (E1, gibberish/special tokens) and semantic adversarial (E2, JailbreakBench/JBB-Behaviors prompts). For each input we record the layer activation, the NLA verbalizer's explanation, the AR reconstruction, raw MSE, and an anisotropy-baseline-corrected (ABC) cosine score following Ethayarajh (2019). Empirical mean pairwise cos of gold activations: Qwen 0.497 vs Gemma 0.975; confirming that raw MSE is not directly comparable across these models. Phase 2: CoT unfaithfulness probe (N=60 per model). 30 factual MCQs in neutral and 3-shot biased conditions. Both models answered correctly on all 60 trials; the Turpin-style unfaithfulness rate is undefined for this setup. Reported as a negative result. Includes all input lists, parquet/JSON results, hallucination judgments under a 4-class rubric and the Python scripts needed to reproduce all experiments from the released NLA checkpoints. License: CC-BY-4.0.



