课题基金 / 基金详情

CAREER: Scalable methods for discovering multivariate dependencies in high dimensional data.

CAREER: Scalable methods for discovering multivariate dependencies in high dimensional data.
职业:用于发现高维数据中多元依赖性的可扩展方法。
批准号:
1352656
负责人:
Balakanapathy Rajaratnam
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-01 至 2019-03-31

项目摘要

项目成果

Balakanapathy Rajaratnam的其他基金

相似基金

相关文献

中文摘要
翻译
这项建议旨在开发原则性的方法来发现多变量依赖关系,以迎合超高维环境。将提出的方法统一起来的一个共同主题是可伸缩性和确定其局限性。识别稀疏逆协方差矩阵的一种流行方法是通过惩罚似然方法。我们提出了一种新的方法来求解惩罚高斯对数似然,其速度比竞争对手快了许多个数量级。该方案中的第二个研究部分研究有限样本中阈值矩阵的统计性质,以期获得具有高度可伸缩性的正定协方差估计方法。该项目的第三个研究方面是对估计的图形网络模型的可变性和不确定性进行量化。介绍了一种利用图解模型选择问题的凸伪似然公式的方法。这允许开发一种具有理论保障的高度可扩展的不确定度量化方法。该项目的第四个研究方面审查了将前三个分部分中提出的方法用于气候变化领域的应用,在气候变化领域,需要进行高维协方差估计。该提案还有一个重要的教学和宣传部分,旨在向有抱负的年轻科学家介绍统计学,这些数据处于本科和研究生学习的不同阶段。来自各种应用的高通量数据的可获得性,包括基因组学、环境科学和其他应用,产生了对分析高维数据的方法和工具的迫切需求。提取和理解数据中的许多复杂关系和多变量依赖关系,并开发有原则的推理程序是统计学家和数据科学家面临的主要挑战之一。本项目中提出的理论和方法工作是在地球和环境科学、基因组学和癌症研究以及社会科学等领域的应用和跨学科合作的推动下进行的。例如,在基因组学中,人们经常感兴趣的是各种基因是如何关联的,以及这些关联在实验(患病)组和对照组之间有何不同。基因调控网络也是研究疾病进化的重要工具。在气候变化辩论的背景下,对全球不同地点的温度进行建模需要对这些变量相关的方式进行简明的建模。建模相关性也自然出现在材料科学和工程中,在这些科学和工程中,人们有兴趣看到当新材料产生时,不同的原子粒子是如何相互作用的。因此,由于理解许多变量之间的关联/关系是许多科学学科共同的努力,因此拟议的在非常高维环境中估计相关性的项目将有广泛的应用。拟议的工作虽然扎根于统计科学,但非常跨学科,涉及统计学家/数据科学家与生物医学科学家、工程师和地球科学家之间的合作和伙伴关系。
英文摘要
This proposal aims to develop principled methods for discovering multivariate dependencies which cater to ultra high dimensional settings. A common theme that unites the proposed methods is scalability and identification of their limitations. A popular approach to identifying sparse inverse covariance matrices is through penalized likelihood methods. We propose a novel approach for solving the penalized Gaussian log-likelihood that is faster than its competitors by many orders of magnitude. The second research component in the proposal investigates the statistical properties of thresholded matrices in finite samples, with a view to obtaining a positive definite covariance estimation method which is highly scalable. The third research aspect of the project investigates quantifying the variability and uncertainty of estimated graphical network models. A methodology that takes advantage of a convex pseudo-likelihood formulation of the graphical model selection problem is introduced. This allows for the development of a highly scalable uncertainty quantification method with theoretical safeguards. The fourth research aspect of the project examines the use of the methodology proposed in the previous three sub-components to an application in the area of climate change, where high dimensional covariance estimation is required. The proposal also has a significant teaching and outreach component which aims to introduce statistics to aspiring young scientists at various stages of their undergraduate and graduate studies.The availability of high-throughput data from various applications, including genomics, environmental sciences and others, has created an urgent need for methodology and tools for analyzing high dimensional data. Extracting and making sense of the many complex relationships and multivariate dependencies in the data and developing principled inferential procedures is one of the major challenges facing statisticians and data scientists. The theoretical and methodological work proposed in this project is motivated by applications and interdisciplinary collaborations in fields as diverse as the earth and environmental sciences, genomics and cancer research, and the social sciences. In genomics for instance, one is often interested to know how various genes are associated, and how these associations differ between an experimental (diseased) and control group. Gene regulatory networks also serve as important tools to study the evolutions of diseases. In the context of the climate change debate, modeling temperature at different points on the globe requires parsimonious modeling of the way in which these variables are related. Modeling correlations also arises naturally in material sciences and engineering where one is interested in seeing how different atomic particles interact when new materials are produced. Hence the proposed project for estimating correlations in very high dimensional settings will have widespread applications, since understanding associations/relationships between many variables is an endeavor that is common to many scientific disciplines. The proposed work, though firmly rooted in the statistical sciences, is very much interdisciplinary, and involves collaborations and partnerships between statisticians/data scientists and biomedical scientists, engineers and earth scientists.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Scalable methods for discovering multivariate dependencies in high dimensional data.
  • 批准号:
    1916787
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $29.32万
  • 财政年份:
    2017
  • 负责人:
    Balakanapathy Rajaratnam
  • 依托单位:
Collaborative Research: Objective Bayesian Model Selection and Estimation in High Dimensional Statistical Models
  • 批准号:
    1106642
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.92万
  • 财政年份:
    2011
  • 负责人:
    Balakanapathy Rajaratnam
  • 依托单位:
CMG Collaborative Research: Efficient high dimensional Bayesian methods for climate field reconstruction
  • 批准号:
    1025465
  • 项目类别:
    Standard Grant
  • 资助金额:
    $35.46万
  • 财政年份:
    2010
  • 负责人:
    Balakanapathy Rajaratnam
  • 依托单位:
Collaborative Research: P2C2--Multiproxy Reconstructions as A Missing-Data Problem: New Techniques and their Application to Regional Climates of the Past Millennium
  • 批准号:
    1003823
  • 项目类别:
    Standard Grant
  • 资助金额:
    $19.79万
  • 财政年份:
    2010
  • 负责人:
    Balakanapathy Rajaratnam
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis