Variable Selection With Second-Generation <i>P</i>-Values
收藏资源简介:
Many statistical methods have been proposed for variable selection in the past century, but few balance inference and prediction tasks well. Here, we report on a novel variable selection approach called penalized regression with second-generation p-values (ProSGPV). It captures the true model at the best rate achieved by current standards, is easy to implement in practice, and often yields the smallest parameter estimation error. The idea is to use an l0 penalization scheme with second-generation p-values (SGPV), instead of traditional ones, to determine which variables remain in a model. The approach yields tangible advantages for balancing support recovery, parameter estimation, and prediction tasks. The ProSGPV algorithm can maintain its good performance even when there is strong collinearity among features or when a high-dimensional feature space with p > n is considered. We present extensive simulations and a real-world application comparing the ProSGPV approach with smoothly clipped absolute deviation (SCAD), adaptive lasso (AL), and minimax concave penalty with penalized linear unbiased selection (MC+). While the last three algorithms are among the current standards for variable selection, ProSGPV has superior inference performance and comparable prediction performance in certain scenarios.
近一个世纪以来,学界已提出诸多用于变量选择的统计方法,但鲜有方法能很好地兼顾统计推断与预测两类任务。本文提出一种新型变量选择方法——基于第二代p值(second-generation p-values, SGPV)的惩罚回归(penalized regression with second-generation p-values, ProSGPV)。该方法能以当前主流方法的最优速率还原真实模型,实际应用中易于实现,且通常能取得最小的参数估计误差。其核心思路为:采用基于第二代p值的L0惩罚框架,而非传统惩罚框架,来筛选模型中保留的变量。该方法在兼顾支撑恢复(support recovery)、参数估计与预测任务方面具备显著优势。即便特征间存在强共线性,或面对p>n的高维特征空间场景,ProSGPV算法仍能保持优异的性能表现。本文通过大量模拟实验与一项真实世界应用案例,将ProSGPV方法与平滑截断绝对偏差(smoothly clipped absolute deviation, SCAD)、自适应Lasso(adaptive lasso, AL)以及带惩罚线性无偏选择的极小极大凹惩罚(minimax concave penalty with penalized linear unbiased selection, MC+)进行了对比。尽管后三种算法均属于当前主流的变量选择方法,但ProSGPV在部分场景下具备更优异的统计推断性能,且预测性能与之相当。




