Group SLOPE – Adaptive Selection of Groups of Predictors
收藏资源简介:
Sorted L-One Penalized Estimation (SLOPE; Bogdan et al. 2013, 2015) is a relatively new convex optimization procedure, which allows for adaptive selection of regressors under sparse high-dimensional designs. Here, we extend the idea of SLOPE to deal with the situation when one aims at selecting whole groups of explanatory variables instead of single regressors. Such groups can be formed by clustering strongly correlated predictors or groups of dummy variables corresponding to different levels of the same qualitative predictor. We formulate the respective convex optimization problem, group SLOPE (gSLOPE), and propose an efficient algorithm for its solution. We also define a notion of the group false discovery rate (gFDR) and provide a choice of the sequence of tuning parameters for gSLOPE so that gFDR is provably controlled at a prespecified level if the groups of variables are orthogonal to each other. Moreover, we prove that the resulting procedure adapts to unknown sparsity and is asymptotically minimax with respect to the estimation of the proportions of variance of the response variable explained by regressors from different groups. We also provide a method for the choice of the regularizing sequence when variables in different groups are not orthogonal but statistically independent and illustrate its good properties with computer simulations. Finally, we illustrate the advantages of gSLOPE in the context of Genome Wide Association Studies. R package grpSLOPE with an implementation of our method is available on The Comprehensive R Archive Network.
排序L1惩罚估计(Sorted L-One Penalized Estimation, SLOPE;Bogdan等人2013、2015)是一类较新型的凸优化方法,可在稀疏高维设计场景下实现回归变量的自适应选择。在此基础上,本文将SLOPE的核心思想拓展至需选择整组解释变量而非单个回归变量的任务场景。此类变量组可通过对强相关预测变量进行聚类得到,或由对应同一定性预测变量不同水平的虚拟变量组构成。本文构建了对应的凸优化问题——分组SLOPE(group SLOPE, gSLOPE),并提出了一种高效的求解算法。此外,本文定义了分组错误发现率(group false discovery rate, gFDR)的概念,并给出了gSLOPE的调参序列选择方案:当变量组之间相互正交时,可证明分组错误发现率被严格控制在预设水平。进一步地,本文证明所提方法可适配未知稀疏性,且在估计由不同组回归变量所解释的响应变量的方差占比时,具有渐近极小极大最优性。本文还提出了一种在不同变量组非正交但统计独立时的正则化序列选择方法,并通过计算机仿真验证了该方法的优良性能。最后,本文通过全基因组关联研究(Genome Wide Association Studies)场景展示了gSLOPE的应用优势。实现本文方法的R包grpSLOPE可在R综合存档网络(The Comprehensive R Archive Network)上获取。



