Automatic Choice of Dimensionality for PCA

Automatic Choice of Dimensionality for PCA
复制标题

DOI:
--
复制
发表时间:
2000
期刊:
--
影响因子:
--
通讯作者:
T. Minka
T. Minka
中科院分区:
其他
文献类型:
--
作者:
T. Minka

文献摘要

被引文献

相似文献

主成分分析(PCA)中的一个中心问题是选择要保留的主成分的数量。通过将PCA解释为密度估计,我们展示了如何使用贝叶斯模型选择来估计数据的真实维度。结果估计计算起来很简单,但保证在有足够数据的情况下选择正确的维度。估计涉及k帧的Steifel流形上的积分,难以精确计算。但在选择合适的参数化并应用拉普拉斯方法后,得到了一个准确实用的估计量。在模拟中,它比交叉验证和其他提出的算法要好得多,而且运行速度快得多。
A central issue in principal component analysis (PCA) is choosing the number of principal components to be retained. By interpreting PCA as density estimation, we show how to use Bayesian model selection to estimate the true dimensionality of the data. The resulting estimate is simple to compute yet guaranteed to pick the correct dimensionality, given enough data. The estimate involves an integral over the Steifel manifold of k-frames, which is difficult to compute exactly. But after choosing an appropriate parameterization and applying Laplace's method, an accurate and practical estimator is obtained. In simulations, it is convincingly better than cross-validation and other proposed algorithms, plus it runs much faster.