The high-dimension, low-sample-size geometric representation holds under mild conditions

The high-dimension, low-sample-size geometric representation holds under mild conditions
复制标题

DOI:
10.1093/biomet/asm050
复制
发表时间:
2007-08-01
期刊:
影响因子:
2.7
通讯作者:
Chi, Yueh-Yun
Chi, Yueh-Yun
中科院分区:
数学2区
文献类型:
--
作者:
Ahn, Jeongyoun;Marron, J. S.;Chi, Yueh-Yun

文献摘要

被引文献

相似文献

高维、小样本数据集具有与传统低维数据不同的几何特性。在他们关于固定样本量增加维数的渐近研究中,Hall等人(2005)表明每个数据向量近似位于高维空间中正则单纯形的顶点上。他们的结果的一个可能不吸引人的方面是潜在的假设,该假设要求被视为时间序列的变量几乎是独立的。利用样本协方差矩阵的渐近性质,我们在更温和的条件下建立了一个等价的几何表示。我们讨论的结果,如使用主成分分析在高维空间,扩展到非独立样本的情况下,也二进制分类问题的影响。
High-dimension, low-small-sample size datasets have different geometrical properties from those of traditional low-dimensional data. In their asymptotic study regarding increasing dimensionality with a fixed sample size, Hall et al. ( 2005) showed that each data vector is approximately located on the vertices of a regular simplex in a high-dimensional space. A perhaps unappealing aspect of their result is the underlying assumption which requires the variables, viewed as a time series, to be almost independent. We establish an equivalent geometric representation under much milder conditions using asymptotic properties of sample covariance matrices. We discuss implications of the results, such as the use of principal component analysis in a high-dimensional space, extension to the case of nonindependent samples and also the binary classification problem.