GPT-EBD2N training logs
收藏资源简介:
Training logs — EBD2N scaling points (IEEE TPDS submission) Four complete training records, one per model size of the scaling table, each the concatenation of the SLURM job logs of the run in chronological order (a ########## job <id> ########## line opens each slice). Every slice resumes exactly where the previous one stopped, at the epoch boundary, with the optimizer and learning-rate schedule restored, so the files show the whole trajectory: configuration, placement, per-epoch throughput and memory, validation accuracy after each epoch, checkpoint writes. ebd2n_254M_training.out — 254 M, MiniPile, 4× H200, 5 epochs, jobs 2249572 + 2255117; validation 47.2 → 52.0 %. ebd2n_483M_training.out — 483 M, 4× H200, jobs 2649838 → 2649840 → 2649839 (two epochs per slice); 54.2 % after 6 epochs (the table stops at epoch 5, same budget as the other sizes). ebd2n_778M_training.out — 778 M, jobs 2656856, 2674201, 2708786, 2718872, 2723380, 2764684 (one epoch per slice); 54.9 % after 6 epochs. ebd2n_1410M_training.out — 1.41 B, 8× H200, pipeline depth 3, jobs 3132735, 3133145, 3133146, 3133147, 3133148 (one epoch per 24 h slice); 49.6 → 55.5 % over 5 epochs. Run dashboard — ebd2n_training_report.html Self-contained HTML report written by the EBD2N runtime at the end of a training run; opens in any browser offline. The file included is the page of one complete run on the CRIANN cluster (job 3242074, 25 September 2026, 4× H200): a one-epoch fine-tuning of Qwen1.5-MoE-A2.7B on text-to-SQL with the first twelve blocks frozen, rank-16 adapters on the trainable blocks, bf16 compute, pipeline depth 3, micro-batch 8 — 9,625 batches, 99.1 % validation token accuracy, 1.87 batches/s, 47.8 GB peak on the most loaded GPU, no skipped optimizer step. It is given as an example of what the runtime records; it is not one of the runs of the scaling table above, which are documented by their logs. The page records, from the training process itself: the run identity (purpose, configuration, host, GPUs, PyTorch/CUDA versions); the hyper-parameters (model, optimisation, numerics, pipeline depth, freeze and LoRA recipe); the metrics (validation accuracy, loss and token accuracy per batch, throughput, allocated and peak memory per GPU, trainable parameters); interactive curves with a table view; the placement of the 216 units over the four GPUs; and, for every unit, its state (frozen by the prefix, frozen as base under LoRA, LoRA, full rank), parameters, optimizer steps taken and skipped. Scores computed afterwards from the saved checkpoint appear in the job log, not on the page. Generated by ebd2n/report.py, included with its test in the code archive.



