遇见数据集

Conditional Sure Independence Screening

收藏
Figshare2018-06-06 更新2026-04-29 收录
官方服务:

资源简介:

Independence screening is powerful for variable selection when the number of variables is massive. Commonly used independence screening methods are based on marginal correlations or its variants. When some prior knowledge on a certain important set of variables is available, a natural assessment on the relative importance of the other predictors is their conditional contributions to the response given the known set of variables. This results in conditional sure independence screening (CSIS). CSIS produces a rich family of alternative screening methods by different choices of the conditioning set and can help reduce the number of false positive and false negative selections when covariates are highly correlated. This article proposes and studies CSIS in generalized linear models. We give conditions under which sure screening is possible and derive an upper bound on the number of selected variables. We also spell out the situation under which CSIS yields model selection consistency and the properties of CSIS when a data-driven conditioning set is used. Moreover, we provide two data-driven methods to select the thresholding parameter of conditional screening. The utility of the procedure is illustrated by simulation studies and analysis of two real datasets. Supplementary materials for this article are available online.

当变量维度庞大时,独立筛选(Independence Screening)是一种极具优势的变量选择手段。当前常用的独立筛选方法多基于边际相关及其变体形式。当已掌握部分重要变量集合的先验知识时,评估其余预测变量相对重要性的自然方式,便是考察其在给定已知变量集合的条件下对响应变量的条件贡献,由此便产生了条件确定独立筛选(CSIS)。通过选取不同的条件集合,CSIS可衍生出丰富的备选筛选方法族;在协变量高度相关的场景下,该方法还能有效降低假阳性与假阴性筛选的数量。本文针对广义线性模型提出并研究了CSIS方法:给出了确保筛选可行的条件,并推导了入选变量数量的上界;同时阐明了CSIS实现模型选择一致性的场景,以及采用数据驱动条件集合时CSIS的相关性质。此外,本文还提出了两种数据驱动的方法,用于选取条件筛选的阈值参数。最后,本文通过仿真实验与两个真实数据集的分析,验证了所提方法的实用性。本文的补充材料可在线获取。

创建时间:
2018-06-06
二维码
社区交流群
二维码
科研交流群
商业服务