遇见数据集

Training and evaluation evidence for "[¬Re] Differential Attention at Small Scale: A Paired, Cross-Stack Reproduction with No Generalization Benefit"

收藏
Zenodo2026-07-25 更新2026-08-01 收录
官方服务:

资源简介:

Raw measurement records behind every number in the article "[¬Re] Differential Attention at Small Scale: A Paired, Cross-Stack Reproduction with No Generalization Benefit". Nothing in the article is hand-typed; tables, figures and macros are generated from these files by docs/paper/gen.py in the code repository (https://github.com/guygrigsby/diff-mlx). Contents: - stage0/: per-step training and held-out validation metrics for the Stage 0 seed band (30M parameters, 100M tokens per run, seeds 1 to 4, each internally paired with byte-identical shared init and identical data order), plus the recorded original seed 0 result with provenance note and the full training logs of the band runs. - stage1/: per-step metrics and run configs for the headline Stage 1 pair (162M parameters, 2.0B tokens, seed 0). Same files as published with the model checkpoints at https://huggingface.co/guygrigsby/diff-mlx. - position_binned_nll.npz: per-token-position held-out negative log-likelihood for both Stage 1 final checkpoints (numpy arrays "diff" and "vanilla", shape (2048,)). Metrics schema (jsonl, one record per logged step): step, train_loss, lr, tps, wall, and periodically val_loss_monitor / val_loss_full (nats, held-out). The original article replicated is Ye et al., Differential Transformer, ICLR 2025 (arXiv:2410.05258). Training ran on one Apple M5 Max in MLX with a PyTorch cross-stack replication on an NVIDIA RTX 3070 Ti.

提供机构:
Zenodo
创建时间:
2026-07-25
二维码
社区交流群
二维码
科研交流群
商业服务