Coded corpus and significance-aware re-benchmark for: Are Tabular Foundation Models Evaluated Fairly? A Four-Axis Critical Survey and Significance-Aware Reporting Standard"
收藏资源简介:
Supplementary data and code for the review submitted to Artificial Intelligence Review. Contents: (1) the full coded corpus of 120 studies scored on the four-axis 0–2 evaluation-rigour rubric (statistical rigour, baseline fairness, data integrity, claim integrity), as XLSX and CSV; (2) the PRISMA-S database search log; (3) the AR-01 significance-aware re-benchmark: full results over 6 models x 25 datasets x 5 seeds (749/750 jobs, one CatBoost-on-mfeat-fourier timeout averaged over 4 seeds) and the computed statistics report; (4) all analysis and figure code (corpus build, job-based benchmark runner, Friedman/Nemenyi/Wilcoxon analysis, critical-difference and axis-profile figures). Data are released under CC-BY-4.0; code under the MIT License (see README).



