High Dimensional Variable Selection with Reciprocal <i>L</i><sub>1</sub>-Regularization
收藏资源简介:
During the past decade, penalized likelihood methods have been widely used in variable selection problems, where the penalty functions are typically symmetric about 0, continuous and nondecreasing in (0, ∞). We propose a new penalized likelihood method, reciprocal Lasso (or in short, rLasso), based on a new class of penalty functions which are decreasing in (0, ∞), discontinuous at 0, and converge to infinity when the coefficients approach zero. The new penalty functions give nearly zero coefficients infinity penalties; in contrast, the conventional penalty functions give nearly zero coefficients nearly zero penalties (e.g., Lasso and SCAD) or constant penalties (e.g., <i>L</i><sub>0</sub> penalty). This distinguishing feature makes rLasso very attractive for variable selection: It can effectively avoid to select overly dense models. We establish the consistency of the rLasso for variable selection and coefficient estimation under both the low and high dimensional settings. Since the rLasso penalty functions induce an objective function with multiple local minima, we also propose an efficient Monte Carlo optimization algorithm to solve the involved minimization problem. Our simulation results show that the rLasso outperforms other popular penalized likelihood methods, such as Lasso, SCAD, MCP, SIS, ISIS and EBIC: It can produce sparser and more accurate coefficient estimates, and catch the true model with a higher probability.
近十年来,惩罚似然方法(penalized likelihood methods)已被广泛应用于变量选择问题(variable selection problems)中,这类方法的惩罚函数通常关于原点对称,且在(0, +∞)区间内连续且非递减。本文提出一种全新的惩罚似然方法——倒数Lasso(reciprocal Lasso,简称rLasso),其基于一类新型惩罚函数:该类函数在(0, +∞)区间内单调递减,在原点处不连续,且当系数趋近于0时趋向于正无穷。此类新型惩罚函数会对趋近于0的系数施加无穷大的惩罚;与之相对,传统惩罚函数则会对趋近于0的系数施加近乎为0的惩罚(如Lasso与SCAD),或施加恒定惩罚(如L₀惩罚(L₀ penalty))。这一显著特性使得rLasso在变量选择任务中极具应用价值:它可有效避免选择过于冗余的模型。本文证明了rLasso在低维和高维场景下,均具备变量选择与系数估计的一致性。由于rLasso的惩罚函数会导致目标函数存在多个局部极小值,本文同时提出一种高效的蒙特卡洛优化算法,以求解该极小化问题。仿真实验结果表明,rLasso的性能优于其他主流惩罚似然方法,包括Lasso、SCAD、MCP、SIS、ISIS与EBIC:它能够生成更稀疏且更精准的系数估计,并以更高概率识别出真实模型。




