Inference in Additively Separable Models With a High-Dimensional Set of Conditioning Variables
收藏资源简介:
This article studies nonparametric series estimation and inference for the effect of a single variable of interest x on an outcome y in the presence of potentially high-dimensional conditioning variables z. The context is an additively separable modelE[y|x,z]=g0(x)+h0(z). The model is high-dimensional in the sense that the series of approximating functions forh0(z)can have more terms than the sample size, thereby allowing z potentially to have very many measured characteristics. The model is required to be approximately sparse:h0(z)can be approximated using only a small subset of series terms whose identities are unknown. This article proposes an estimation and inference method forg0(x)called Post-Nonparametric Double Selection, which is a generalization of Post-Double Selection. Rates of convergence and asymptotic normality for the estimator are derived and hold over a large class of sparse data-generating processes. A simulation study illustrates finite sample estimation properties of the proposed estimator and coverage properties of the corresponding confidence intervals. Finally, an empirical application to college admissions policy demonstrates the practical implementation of the proposed method.
本文针对存在潜在高维条件变量z的场景,围绕单一关注变量x对结果变量y的效应展开非参数序列估计(nonparametric series estimation)与统计推断研究。本文采用可加分离模型框架:条件期望E[y|x,z] = g0(x) + h0(z)。该模型属于高维设定:针对h0(z)的逼近函数序列所含项数可超过样本量,从而允许z包含极多可测特征。该模型满足近似稀疏性要求:仅需利用身份未知的少量序列项子集,即可对h0(z)实现有效逼近。本文针对g0(x)提出一种名为后非参数双重选择(Post-Nonparametric Double Selection)的估计与推断方法,该方法是后双重选择(Post-Double Selection)的推广形式。本文推导了该估计量的收敛速率与渐近正态性,且该性质在一大类稀疏数据生成过程中均成立。本文通过模拟实验验证了所提估计量的有限样本估计性能,以及对应置信区间的覆盖性能。最后,本文将所提方法应用于大学招生政策的实证分析,展示了该方法的实际实施路径。



