Reproducibility Data: Subspace Direction, Not Rank - Scaling Linear Attention via Sweet-Spot Init Transfer
收藏资源简介:
Experimental data underlying the paper 'Subspace Direction, Not Rank: Scaling Linear Attention via Sweet-Spot Init Transfer'. Contains structured per-seed results (best/final accuracy + training curves), SVD-based effective-rank and cosine-similarity analyses (subspace direction alignment), five-task probing comparisons, MQAR benchmark results, direction-ablation results, and raw training logs. Hardware: NVIDIA V100 and MetaX C550 (8-GPU DDP). The included README maps every table and figure in the paper to its source data file, enabling independent verification.What is new in this version: (1) the four d4096_seq24_* files document the within-protocol (seq=24) re-test of d=4096 at its SNR-matched lr (5e-4) reported in §3.1 as a negative result (does not break within 100k, unlike the companion-study seq=64 configuration); (2) results_mqar_raw.json preserves the raw d=2048 MQAR accuracy dictionaries cited in Table 7; (3) the four causal_* files support the self-reinforcing phase-transition discussion in §6 (gradient-amplitude intervention on breaking and stuck configurations); (4) results_mqar_full.md is expanded with the key findings and direction-ablation summary. v3 (2026-08-09): adds rebuilt_chain/ with the pre-registered end-to-end chain reconstruction evidence (README + final logs for 4 seeds: 0, 1, 42, 123; all seeds reach LRR = 1.000 at d=4096), reported in the paper's Section 5.5.



