PROJECTED PRINCIPAL COMPONENT ANALYSIS IN FACTOR MODELS.

PROJECTED PRINCIPAL COMPONENT ANALYSIS IN FACTOR MODELS.
复制标题

在因子模型中预计主成分分析。

DOI:
10.1214/15-aos1364
复制
发表时间:
2016-02
影响因子:
4.5
通讯作者:
Wang W
Wang W
中科院分区:
数学1区
文献类型:
--
作者:
Fan J;Liao Y;Wang W

文献摘要

被引文献

相似文献

本文介绍了一种投影主成分分析(Projected-PCA),它将主成分分析投影(平滑)到由协变量构成的给定线性空间上的数据矩阵。当它应用于高维因子分析时,投影去除了噪声分量。我们表明,未观察到的潜在因素可以更准确地估计比传统的PCA,如果投影是真实的,或更精确地说,当因子加载矩阵与投影线性空间。当维数较大时,即使在样本容量有限的情况下,也可以准确地估计因子。我们提出了一个灵活的半参数因子模型,该模型将因子载荷矩阵分解为可以由特定主题的协变量和正交残差分量解释的分量。协变量对因子负荷的影响进一步通过加性模型通过筛选近似来建模。利用新提出的投影主成分分析方法,得到了光滑因子加载矩阵的收敛速度,比传统的因子分析方法快得多。即使在样本量有限的情况下也能实现收敛,并且在高维低样本量的情况下特别有吸引力。这导致我们开发非参数检验是否观察到的协变量有解释力的负载,他们是否完全解释了负载。模拟数据和S&P 500指数成分股的收益率都说明了所提出的方法。
This paper introduces a Projected Principal Component Analysis (Projected-PCA), which employees principal component analysis to the projected (smoothed) data matrix onto a given linear space spanned by covariates. When it applies to high-dimensional factor analysis, the projection removes noise components. We show that the unobserved latent factors can be more accurately estimated than the conventional PCA if the projection is genuine, or more precisely, when the factor loading matrices are related to the projected linear space. When the dimensionality is large, the factors can be estimated accurately even when the sample size is finite. We propose a flexible semi-parametric factor model, which decomposes the factor loading matrix into the component that can be explained by subject-specific covariates and the orthogonal residual component. The covariates’ effects on the factor loadings are further modeled by the additive model via sieve approximations. By using the newly proposed Projected-PCA, the rates of convergence of the smooth factor loading matrices are obtained, which are much faster than those of the conventional factor analysis. The convergence is achieved even when the sample size is finite and is particularly appealing in the high-dimension-low-sample-size situation. This leads us to developing nonparametric tests on whether observed covariates have explaining powers on the loadings and whether they fully explain the loadings. The proposed method is illustrated by both simulated data and the returns of the components of the S&P 500 index.