课题基金 / 基金详情

Discovering Sparse Covariance Structures in High Dimensions

Discovering Sparse Covariance Structures in High Dimensions
发现高维稀疏协方差结构
批准号:
0805798
负责人:
Elizaveta Levina
金额:
$25.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-06-01 至 2012-05-31

项目摘要

项目成果

Elizaveta Levina的其他基金

相似基金

相关文献

中文摘要
翻译
这个项目的重点是发现和利用数据中的稀疏结构来改进高维协方差矩阵的估计。协方差矩阵在许多数据分析方法中起着关键作用,包括主成分分析、判别分析、多元分析中均值的推断以及图模型中独立性和条件独立性关系的推断。随机矩阵理论的进展表明,传统的估计方法,即样本协方差,在高维时表现不佳。现有的关于替代估计的研究,包括PI的前人的工作,大多集中在变量指标(时间序列、纵向数据、空间数据、光谱等)存在距离或排序的情况下。然而,在许多应用中,这种排序是不可用的:例如,遗传学、金融、社会和经济数据。这个项目开发了几种构造正则化稀疏估计器的方法,这些估计器对变量排列不变,对协方差矩阵及其逆都是如此。这些方法的主要组成部分是阈值、鼓励稀疏性的平滑惩罚、排列不变的损失函数、自适应权重和发现变量潜在结构化重新排序的多重投影。充分发展了所提出的高维估计的一致性和收敛速度的分析结果。这些高维的理论结果需要不同于标准渐近分析的工具,现有文献中几乎没有可用的工具。开发了计算这些估计量所需的高效优化算法,重点是计算成本随着维度的增加而尽可能缓慢地增长。其中一些估计器的设计计算成本非常低,而另一些估计器则需要计算的独创性才能在真正高维的情况下可行。所提出的方法在模拟和通过PI的跨学科合作的许多应用中都得到了广泛的测试。在现代世界中收集的大量数据给统计学家带来了新的挑战。迫切需要处理高维数据的新的理论和实践方法,以及需要估计高维协方差矩阵作为数据分析的一部分的大量应用:金融、遗传学、光谱学、遥感、气候研究、脑成像、语音识别等。PI正在与化学家就骨骼的拉曼光谱进行合作,与海洋学家就将光谱数据用于遥感海洋进行合作,与气候科学家就温度建模进行合作,并与生物统计学家就一种在蛋白质水平上工作的新型基因表达技术进行合作。PI还活跃在无线传感器网络的统计信号处理领域,其中空间协方差估计是重要的,并且具有许多安全应用。本项目开发的高维协方差估计新方法在理论上进行了分析,并在这些应用程序中进行了测试和验证,反过来,项目在后期阶段的发展方向受到应用程序的问题和需求的影响。该项目还有助于培养现代统计学一个重要领域的研究生。
英文摘要
This project focuses on discovering and exploiting sparse structures in the data to improve estimation of covariance matrices in high dimensions. The covariance matrix plays a key role in many data analysis methods, including principal component analysis, discriminant analysis, inference about the means in multivariate analysis, and inference about independence and conditional independence relationships in graphical models. Advances in random matrix theory have shown that the traditional estimator, the sample covariance, performs poorly in high dimensions. The existing research on alternative estimators, including previous work of the PI, focuses mostly on the situation when there is a notion of distance or ordering for the variable indexes (time series, longitudinal data, spatial data, spectroscopy, etc). However, there are many applications where such ordering is not available: for example, genetics, financial, social and economic data. This project develops several methods for constructing regularized sparse estimators that are invariant to variable permutations, both for the covariance matrix and its inverse. The main building blocks of the methods are thresholding, smooth penalties that encourage sparsity, permutation-invariant loss functions, adaptive weights, and manifold projections to discover potential structured re-orderings of the variables. Analytical results establishing consistency and convergence rates of the proposed estimators in high dimensions are fully developed. These theoretical results in high dimensions require tools that are different from standard asymptotic analysis, and there are few available in the existing literature. Efficient optimization algorithms needed to compute these estimators are developed, with the emphasis on the computational cost growing as slowly as possible with dimension. Some of the estimators proposed carry a very low computation cost by design, while others require computational ingenuity to be feasible in really high dimensions. The proposed methodology is tested extensively, both in simulations and on a number of applications through the PI's interdisciplinary collaborations.Massive amounts of data collected in the modern world are creating new challenges for statisticians. There is an urgent need for new theoretical and practical methods that deal with high-dimensional data, and a vast number of applications where high-dimensional covariance matrices need to be estimated as part of data analysis: finance, genetics, spectroscopy, remote sensing, climate studies, brain imaging, speech recognition, and many others. The PI has ongoing collaborations with chemists on Raman spectroscopy of bone, with oceanologists on using spectral data for remote ocean sensing, with climate scientists on temperature modeling and with a biostatistician on a new type of gene expression technology that works at protein level. The PI also works actively in the area of statistical signal processing by wireless sensor networks, where spatial covariance estimation is important, and which has many security applications. The new methodology for estimating high-dimensional covariances developed in this project is analyzed theoretically and tested and validated in these applications, and in turn, the directions in which the project develops at later stages are influenced by the issues and needs of the applications. The project also contributes to educating graduate students in an important area of modern statistics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FRG: Collaborative Research: Flexible Network Inference
Multivariate Analysis for Samples of Networks
RTG: Understanding dynamic big data with complex structure
Conference proposal: From Industrial Statistics to Data Science
国内基金
海外基金
基于Sparse-Land模型的SAR图像噪声抑制与分割
  • 批准号:
    60971128
  • 项目类别:
    面上项目
  • 资助金额:
    30.0万元
  • 批准年份:
    2009
  • 负责人:
    侯彪
  • 依托单位: