遇见数据集

KAN-Payne: data package for a controlled evaluation of Kolmogorov–Arnold network spectral emulators (version 2)

收藏
Zenodo2026-08-07 更新2026-08-13 收录
官方服务:

资源简介:

What is here Every table and figure in the paper is reproducible from these records without retraining. The trained checkpoints are included, so the emulators can also be re-evaluated or redeployed directly. | Directory | Files | Size | Contents ||---|---:|---:|---|| `01_grids/` | 10 | 65 MiB | Frozen per-seed 720/80/200 grids, and the contract assertions verifying that wavelengths, masks, targets and label bounds are byte-identical to the public release. || `02_tuning/` | 28 | 21 MiB | Per-architecture learning-rate × weight-decay tuning, scored on the 80-star validation split only. || `03_benchmark_runs/` | 116 | 518 MiB | The twelve budget-matched formal runs (4 architectures × 3 seeds): training histories, checkpoints, validation and test metrics, and test-set residual arrays. || `04_cost_benchmark/` | 20 | 40 KiB | Training wall-clock, peak GPU memory, inference latency, Jacobian and per-star fit timing (Table 1). || `05_pretrained_audit/` | 6 | 62 MiB | Audit of the public pretrained Payne network on the same split, with its train- and test-set residual arrays. || `06_injection_recovery/` | 185 | 129 MiB | Known-truth injection–recovery: the 200 frozen test spectra injected at S/N = ∞, 100, 50, 20 and refit for all 25 labels. || `07_multistart/` | 66 | 11 MiB | Randomised multistart: the stratified 200-star APOGEE subset and the synthetic 50-star arm. || `08_data_scaling/` | 42 | 226 KiB | Training-set-size axis, *N* ∈ {100, 200, 400, 720}, three seeds, 60 000-update half budget. || `09_loss_ablation/` | 8 | 16 MiB | L1-versus-MSE training-loss ablation for KAN-L. || `10_supplementary_controls/` | 405 | 778 MiB | Large-MLP shape controls, weight-decay sweeps, the new-protocol APOGEE rerun, the three-seed recovery grid, and the activation and spline-grid-range controls, with their independent re-verification. || `11_analysis/` | 257 | 71 MiB | Everything computed *from* the run records: the numbers behind each table, figure and quoted statistic, in four stages. || `12_tables_and_figures/` | 30 | 3 MiB | The six generated LaTeX tables, `PROVENANCE.json` mapping every printed value to its source record, and the manuscript figures. || `13_provenance_and_audit/` | 36 | 12 MiB | Command log, environment capture, code state, source checksums, artifact inventory, and the independent audit: recomputed metrics, numerical-consistency and reproducibility reports, and the test-set leakage audit. || `code/` | 40 | 1 MiB | The analysis, table and figure generators (identical to `analysis/` in the code repository). | Totals and per-file checksums are in `SHA256SUMS.txt` and `MANIFEST.json`.`PATH_MAP.md` maps the published directory names onto the run labels that appear inside the archived records. Verification status All files carry SHA256 checksums (`SHA256SUMS.txt`). Every metric reported in the paper — 142 of 142 — was independently recomputed from the archived arrays and matches. Test-set access controls were independently confirmed: hyperparameter selection and checkpoint selection recorded zero test-set evaluations, and every test evaluation postdates the selection freeze (`13_provenance_and_audit/final_audit/test_leakage_audit.md`).

提供机构:
Zenodo
创建时间:
2026-08-07
二维码
社区交流群
二维码
科研交流群
商业服务