遇见数据集

Canonical attention data for four open-weight LLMs (Qwen3-4B/8B, Llama-3.1-8B, Mistral-7B) on Chinese and English inputs

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

This dataset accompanies the preprint Attention Convergence Across Architectures: Cross-Lingual Structure and Surrogate-Dependent Causal Effects (Wan-Jiu Yang). Contents. One JSON file per model–language configuration (8 in total): consistency_<model>_<lang>.json for Qwen3-4B, Qwen3-8B, Llama-3.1-8B, and Mistral-7B, each with Chinese (zh) and English (en) inputs. Each file stores the per-layer, per-head attention matrices recorded on a single canonical long-form input (252–458 tokens; causal masking, no chat template, fp16), together with the derived statistics used in the paper: normalized attention entropy (per-row and support-normalized mean H̄), consecutive-layer consistency C_consec(ℓ) and final-layer consistency C_final(ℓ), the position-independence ratio ρ_pos with 25/50/25 depth-band means, per-position column-mass profiles, and convergence/compression phase statistics. The 8,704 layer–head units (136 model–layers × 32 heads × 2 languages) underlie Tables 1, 2, 6, 7 and the observational figures of the paper. Companion repository. All code, the intervention study for eleven model configurations (four attention-replacement modes plus a matched-global control), per-probe and per-sample streaming logs, probe datasets with per-model filtering metadata, and evaluation scripts are available at https://github.com/beta-going/attention-convergence-code. Files. 8 JSON files, ~2.4 GB total.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务