Data for: Training data provenance, not architecture, is the primary determinant of performance on a materials discovery benchmark
收藏数据链接:
官方服务:
资源简介:
This dataset includes per-material predictions, variance decomposition results, error-correlation analyses, scaling analyses, and collective failure characterisations for 45 models (with 53-model sensitivity checks) on the Matbench Discovery benchmark covering 256,963 WBM structures. The analysis demonstrates that training-data provenance explains substantially more performance variance than architecture choice (partial η² of 0.84–0.89 for data versus 0.15–0.35 for architecture), a finding robust under family-aware resampling and label-permutation tests.
创建时间:
2026-03-30



