Effective PCA for high-dimension, low-sample-size data with singular value decomposition of cross data matrix

Effective PCA for high-dimension, low-sample-size data with singular value decomposition of cross data matrix
复制标题

DOI:
10.1016/j.jmva.2010.04.006
复制
发表时间:
2010-10
期刊:
J. Multivar. Anal.
影响因子:
--
通讯作者:
K. Yata;M. Aoshima
K. Yata;M. Aoshima
中科院分区:
其他
文献类型:
--
作者:
K. Yata;M. Aoshima

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的方法来处理PCA在高维,低样本量(HDLSS)数据的情况下。给出了利用交叉数据矩阵的奇异值估计特征值的方法。当维数d和样本量n都以n远小于d的方式增长到无穷大时,我们给出了特征值估计及其极限分布的一致性性质.我们应用新的方法来估计PC方向和PC分数在HDLSS数据的情况下。我们将本文的研究结果应用于混合模型,将数据集分类为两个聚类。我们展示了如何使用HDLSS数据的前列腺癌的微阵列研究的新方法进行。
In this paper, we propose a new methodology to deal with PCA in high-dimension, low-sample-size (HDLSS) data situations. We give an idea of estimating eigenvalues via singular values of a cross data matrix. We provide consistency properties of the eigenvalue estimation as well as its limiting distribution when the dimension d and the sample size n both grow to infinity in such a way that n is much lower than d. We apply the new methodology to estimating PC directions and PC scores in HDLSS data situations. We give an application of the findings in this paper to a mixture model to classify a dataset into two clusters. We demonstrate how the new methodology performs by using HDLSS data from a microarray study of prostate cancer.