Canonical attention data for four open-weight LLMs (Qwen3-4B/8B, Llama-3.1-8B, Mistral-7B) on Chinese and English inputs
收藏资源简介:
This dataset accompanies the preprint Attention Convergence Across Architectures: Cross-Lingual Structure and Surrogate-Dependent Causal Effects (Wan-Jiu Yang). Contents. One JSON file per model–language configuration (8 in total): consistency_<model>_<lang>.json for Qwen3-4B, Qwen3-8B, Llama-3.1-8B, and Mistral-7B, each with Chinese (zh) and English (en) inputs. Each file stores the per-layer, per-head attention matrices recorded on a single canonical long-form input (252–458 tokens; causal masking, no chat template, fp16), together with the derived statistics used in the paper: normalized attention entropy (per-row and support-normalized mean H̄), consecutive-layer consistency C_consec(ℓ) and final-layer consistency C_final(ℓ), the position-independence ratio ρ_pos with 25/50/25 depth-band means, per-position column-mass profiles, and convergence/compression phase statistics. The 8,704 layer–head units (136 model–layers × 32 heads × 2 languages) underlie Tables 1, 2, 6, 7 and the observational figures of the paper. Companion repository. All code, the intervention study for eleven model configurations (four attention-replacement modes plus a matched-global control), per-probe and per-sample streaming logs, probe datasets with per-model filtering metadata, and evaluation scripts are available at https://github.com/beta-going/attention-convergence-code. Files. 8 JSON files, ~2.4 GB total.



