A Representation–Hypothesis–Estimation Decomposition of Prediction Error
收藏资源简介:
Modern predictive modeling provides many ways to evaluate model perfor- mance but few principled tools for diagnosing why a model fails. We introduce the Representation–Hypothesis–Estimation (RHE) decomposition, which par- titions the excess risk into three orthogonal components: representation loss Lrep, measuring information discarded by the representation map; hypothesis- class loss Lhyp, measuring approximation error of the hypothesis class act- ing on the retained representation; and estimation loss Lest, measuring finite- sample deviation from the best achievable predictor. The decomposition fol- lows from a single Pythagorean identity in L2(PX) and applies, under fixed- representation and closed-linear-hypothesis-class conditions, to ordinary least squares, ridge regression, principal component regression, generalized additive models, and fixed-partition CART. We introduce a normalized diagnostic index ρ = (ρrep, ρhyp, ρest) ∈ ∆2, establish joint consistency of the cross-fitted plug-in estimator, and derive a non-asymptotic concentration bound giving the sample size sufficient for reliable diagnosis. A nested cross-validation experiment across 17 datasets (66 comparisons) shows the RHE-recommended intervention outper- forms uninformed model substitution in 71% of comparisons overall and in 95% where the hypothesis class is the dominant bottleneck, with a median held-out MSE reduction of 46% versus −19% for the uninformed alternative.



