Statistical challenges of high-dimensional data INTRODUCTION

Statistical challenges of high-dimensional data INTRODUCTION
复制标题

DOI:
10.1098/rsta.2009.0159
复制
发表时间:
2009-11-13
影响因子:
5
通讯作者:
Titterington, D. Michael
Titterington, D. Michael
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Johnstone, Iain M.;Titterington, D. Michael

文献摘要

被引文献

相似文献

统计理论和方法的现代应用可能涉及极其庞大的数据集,通常在相对较少的实验单元上进行大量测量。作为回应,出现了新的方法和相关理论:本期主题刊的目的是说明这些最新发展中的一些。这篇概括性的文章介绍了在非常熟悉的线性统计模型的背景下使用高维数据时出现的困难:我们给出了当感兴趣的参数向量是稀疏的,即包含许多零元素时可以实现的体验。我们描述了识别包含所有有用信息的数据空间的低维子空间的其他方法。然后回顾分类的主题以及从一个非常大的集合中识别有助于对观察进行分类的变量的问题。简要介绍了高维数据的可视化和处理贝叶斯分析中计算问题的方法。在适当的地方,参考了本期的其他论文。
Modern applications of statistical theory and methods can involve extremely large datasets, often with huge numbers of measurements on each of a comparatively small number of experimental units. New methodology and accompanying theory have emerged in response: the goal of this Theme Issue is to illustrate a number of these recent developments. This overview article introduces the difficulties that arise with high-dimensional data in the context of the very familiar linear statistical model: we give a taste of what can nevertheless be achieved when the parameter vector of interest is sparse, that is, contains many zero elements. We describe other ways of identifying low-dimensional subspaces of the data space that contain all useful information. The topic of classification is then reviewed along with the problem of identifying, from within a very large set, the variables that help to classify observations. Brief mention is made of the visualization of high-dimensional data and ways to handle computational problems in Bayesian analysis are described. At appropriate points, reference is made to the other papers in the issue.