Geometric consistency of principal component scores for high‐dimensional mixture models and its application

Geometric consistency of principal component scores for high‐dimensional mixture models and its application
复制标题

DOI:
10.1111/sjos.12432
复制
发表时间:
2019-12
影响因子:
1
通讯作者:
K. Yata;M. Aoshima
K. Yata;M. Aoshima
中科院分区:
数学4区
文献类型:
--
作者:
K. Yata;M. Aoshima

文献摘要

相似文献

在本文中,我们考虑基于高维混合模型的主成分分析(PCA)的聚类。我们提出了 PCA 对于高维数据聚类有效的理论原因。首先,我们推导出从二类混合模型中获取的高维、低样本量(HDLSS)数据的几何表示。借助几何表示,我们给出了 HDLSS 背景下样本主成分分数的几何一致性属性。我们开发了几何表示的思想,并为多类混合模型提供了几何一致性属性。我们证明 PCA 可以在某些条件下以一种令人惊讶的明确方式对 HDLSS 数据进行聚类。最后,我们展示了使用基因表达数据集进行聚类的性能。
In this article, we consider clustering based on principal component analysis (PCA) for high‐dimensional mixture models. We present theoretical reasons why PCA is effective for clustering high‐dimensional data. First, we derive a geometric representation of high‐dimension, low‐sample‐size (HDLSS) data taken from a two‐class mixture model. With the help of the geometric representation, we give geometric consistency properties of sample principal component scores in the HDLSS context. We develop ideas of the geometric representation and provide geometric consistency properties for multiclass mixture models. We show that PCA can cluster HDLSS data under certain conditions in a surprisingly explicit way. Finally, we demonstrate the performance of the clustering using gene expression datasets.