High dimensional data analysis using multivariate generalized spatial quantiles.

High dimensional data analysis using multivariate generalized spatial quantiles.
复制标题

DOI:
10.1016/j.jmva.2010.12.002
复制
发表时间:
2011-04
影响因子:
1.6
通讯作者:
Chatterjee S
Chatterjee S
中科院分区:
数学2区
文献类型:
--
作者:
Mukhopadhyay ND;Chatterjee S

文献摘要

相似文献

高维数据经常出现在图像分析、遗传实验、网络分析和各种其他研究领域。许多这样的数据集并不对应于研究得很好的概率分布,并且在一些应用中,数据云突出地显示出非对称和非凸形特征。我们建议使用空间分位数及其推广,特别是投影分位数,来描述、分析和进行多变量数据的推理。我们对潜在概率分布的性质和形状特征做了最小的假设,并且我们不要求样本量与数据维度一样高。我们给出了广义空间分位数的理论性质,并给出了一个快速计算它们的算法。我们的分位数可以用来获得不需要符合预定形状的多维置信度或可信区域。我们还提出了一种新的多维顺序统计量的概念,它可以用来获得多维异常值。如果数据被硬塞进众所周知的概率配置中,使用基于广义空间分位数的分析揭示的许多特征将被遗漏。
High dimensional data routinely arises in image analysis, genetic experiments, network analysis, and various other research areas. Many such datasets do not correspond to well-studied probability distributions, and in several applications the data-cloud prominently displays non-symmetric and non-convex shape features. We propose using spatial quantiles and their generalizations, in particular, the projection quantile, for describing, analyzing and conducting inference with multivariate data. Minimal assumptions are made about the nature and shape characteristics of the underlying probability distribution, and we do not require the sample size to be as high as the data-dimension. We present theoretical properties of the generalized spatial quantiles, and an algorithm to compute them quickly. Our quantiles may be used to obtain multidimensional confidence or credible regions that are not required to conform to a pre-determined shape. We also propose a new notion of multidimensional order statistics, which may be used to obtain multidimensional outliers. Many of the features revealed using a generalized spatial quantile-based analysis would be missed if the data was shoehorned into a well-known probabilistic configuration.