PERMUTATION METHODS FOR FACTOR ANALYSIS AND PCA

PERMUTATION METHODS FOR FACTOR ANALYSIS AND PCA
复制标题

DOI:
10.1214/19-aos1907
复制
发表时间:
2020-10-01
影响因子:
4.5
通讯作者:
Dobriban, Edgar
Dobriban, Edgar
中科院分区:
数学1区
文献类型:
--
作者:
Dobriban, Edgar

文献摘要

被引文献

相似文献

研究人员通常有测量样本特征的数据集,比如学生的考试成绩。在因子分析和PCA中,这些特征被认为受到未观察到的因素的影响,例如技能。我们能确定有多少组件影响数据吗?这是一个重要的问题,因为在这里做出的决策对所有下游数据分析都有很大的影响。因此,开发了许多方法。并行分析是一种流行的排列方法:它随机打乱数据的每个特征。如果组件的奇异值大于排列数据的奇异值,则选择这些组件。尽管被广泛使用,也有经验证据证明其准确性,但目前尚无理论依据。在本文中,我们证明了并行分析(或排列方法)一致地选择了某些高维因子模型中的大成分。但是,当信号太大时,不选择较小的组件。直觉是,排列保持噪声不变,同时“破坏”低秩信号。这为排列方法提供了理由。我们的工作也揭示了排列方法的缺点,并为改进铺平了道路。
Researchers often have datasets measuring features xij of samples, such as test scores of students. In factor analysis and PCA, these features are thought to be influenced by unobserved factors, such as skills. Can we determine how many components affect the data? This is an important problem, because decisions made here have a large impact on all downstream data analysis. Consequently, many approaches have been developed. Parallel Analysis is a popular permutation method: it randomly scrambles each feature of the data. It selects components if their singular values are larger than those of the permuted data. Despite widespread use, as well as empirical evidence for its accuracy, it currently has no theoretical justification.In this paper, we show that parallel analysis (or permutation methods) consistently select the large components in certain high-dimensional factor models. However, when the signals are too large, the smaller components are not selected. The intuition is that permutations keep the noise invariant, while "destroying" the low-rank signal. This provides justification for permutation methods. Our work also uncovers drawbacks of permutation methods, and paves the way to improvements.