Inference in Sparsity-Induced Weak Factor Models
收藏资源简介:
In this paper, we consider statistical inference for high-dimensional approximate factor models. We posit a weak factor structure, in which the factor loading matrix can be sparse and the signal eigenvalues may diverge more slowly than the cross-sectional dimension, N. We propose a novel inferential procedure to decide whether each component of the factor loadings is zero or not, and prove that this controls the false discovery rate (FDR) below a pre-assigned level, while the power tends to unity. This “factor selection” procedure is primarily based on a debiased version of the SOFAR estimator of Uematsu and Yamagata (2021), but is also applicable to the principal component (PC) estimator. After the factor selection, the re-sparsified SOFAR and sparsified PC estimators are proposed and their consistency is established. Finite sample evidence supports the theoretical results. We apply our method to the FRED-MD dataset of macroeconomic variables and the monthly firm-level excess returns which constitute the S&P 500 index. The results give very strong statistical evidence of sparse factor loadings under the identification restrictions and exhibit clear associations of factors and categories of the variables. Furthermore, our method uncovers a very weak but statistically significant factor in the residuals of Fama-French five factor regression.
本文针对高维近似因子模型的统计推断问题展开研究。我们设定弱因子结构,其中因子载荷矩阵可为稀疏矩阵,且信号特征值的发散速度慢于横截面维度N。我们提出了一种全新的推断程序,用于检验因子载荷的各分量是否为零,并证明该程序可将错误发现率(false discovery rate, FDR)控制在预设水平以下,同时检验功效趋近于1。该“因子选择”程序主要基于Uematsu与Yamagata(2021)提出的SOFAR估计量的去偏形式,同时也适用于主成分(principal component, PC)估计量。在完成因子选择后,我们进一步提出了再稀疏化SOFAR估计量与稀疏化PC估计量,并证明了二者的一致性。有限样本实证结果验证了理论结论的有效性。我们将所提方法应用于宏观经济变量的FRED-MD数据集,以及构成标准普尔500指数的月度企业层面超额收益率数据。实证结果为识别约束下稀疏因子载荷的存在性提供了强有力的统计证据,同时清晰展现了因子与各类变量间的关联关系。此外,我们的方法还在法玛-弗伦奇五因子回归的残差项中,发现了一个强度极弱但统计显著的因子。




