遇见数据集

A Representation–Hypothesis–Estimation Decomposition of Prediction Error

收藏
Zenodo2026-07-05 更新2026-08-01 收录
官方服务:

资源简介:

Modern predictive modeling provides many ways to evaluate model perfor- mance but few principled tools for diagnosing why a model fails. We introduce the Representation–Hypothesis–Estimation (RHE) decomposition, which par- titions the excess risk into three orthogonal components: representation loss Lrep, measuring information discarded by the representation map; hypothesis- class loss Lhyp, measuring approximation error of the hypothesis class act- ing on the retained representation; and estimation loss Lest, measuring finite- sample deviation from the best achievable predictor. The decomposition fol- lows from a single Pythagorean identity in L2(PX) and applies, under fixed- representation and closed-linear-hypothesis-class conditions, to ordinary least squares, ridge regression, principal component regression, generalized additive models, and fixed-partition CART. We introduce a normalized diagnostic index ρ = (ρrep, ρhyp, ρest) ∈ ∆2, establish joint consistency of the cross-fitted plug-in estimator, and derive a non-asymptotic concentration bound giving the sample size sufficient for reliable diagnosis. A nested cross-validation experiment across 17 datasets (66 comparisons) shows the RHE-recommended intervention outper- forms uninformed model substitution in 71% of comparisons overall and in 95% where the hypothesis class is the dominant bottleneck, with a median held-out MSE reduction of 46% versus −19% for the uninformed alternative.

提供机构:
Zenodo
创建时间:
2026-07-05
二维码
社区交流群
二维码
科研交流群
商业服务