CAREER: Scalable methods for discovering multivariate dependencies in high dimensional data.
CAREER: Scalable methods for discovering multivariate dependencies in high dimensional data.
批准号:
1352656
负责人:
Balakanapathy Rajaratnam
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-01 至 2019-03-31
中文摘要
该提案旨在开发发现多维依赖关系的原则方法,以满足超高维设置。将所提出的方法联合起来的一个共同主题是可伸缩性和对其局限性的识别。一种常用的识别稀疏逆协方差矩阵的方法是惩罚似然方法。我们提出了一种新的方法来解决惩罚高斯对数似然,比其竞争对手快了许多个数量级。第二部分研究了有限样本中阈值矩阵的统计性质,以期得到一种高度可扩展的正定协方差估计方法。该项目的第三个研究方面是对估计的图形网络模型的可变性和不确定性进行量化。介绍了一种利用凸伪似然公式求解图形模型选择问题的方法。这允许开发具有理论保障的高度可扩展的不确定性量化方法。该项目的第四个研究方面考察了前三个子部分中提出的方法在气候变化领域的应用,该领域需要高维协方差估计。该提案还有一个重要的教学和推广部分,旨在向处于本科和研究生学习不同阶段的有抱负的年轻科学家介绍统计学。从基因组学、环境科学和其他应用中获得的高通量数据产生了对分析高维数据的方法和工具的迫切需求。提取和理解数据中的许多复杂关系和多元依赖关系,并开发有原则的推理程序是统计学家和数据科学家面临的主要挑战之一。在这个项目中提出的理论和方法工作是由应用和跨学科合作的领域,如地球和环境科学,基因组学和癌症研究,以及社会科学。例如,在基因组学中,人们通常有兴趣知道各种基因是如何关联的,以及这些关联在实验组(患病组)和对照组之间有何不同。基因调控网络也是研究疾病进化的重要工具。在气候变化辩论的背景下,对地球上不同地点的温度进行建模需要对这些变量之间的关联方式进行简洁的建模。在材料科学和工程中,建模相关性也很自然地出现,人们对观察不同的原子粒子在生产新材料时如何相互作用感兴趣。因此,在非常高维环境中估计相关性的提议项目将具有广泛的应用,因为理解许多变量之间的关联/关系是许多科学学科共同的努力。拟议的工作虽然牢牢扎根于统计科学,但在很大程度上是跨学科的,涉及统计学家/数据科学家与生物医学科学家、工程师和地球科学家之间的合作和伙伴关系。
英文摘要
This proposal aims to develop principled methods for discovering multivariate dependencies which cater to ultra high dimensional settings. A common theme that unites the proposed methods is scalability and identification of their limitations. A popular approach to identifying sparse inverse covariance matrices is through penalized likelihood methods. We propose a novel approach for solving the penalized Gaussian log-likelihood that is faster than its competitors by many orders of magnitude. The second research component in the proposal investigates the statistical properties of thresholded matrices in finite samples, with a view to obtaining a positive definite covariance estimation method which is highly scalable. The third research aspect of the project investigates quantifying the variability and uncertainty of estimated graphical network models. A methodology that takes advantage of a convex pseudo-likelihood formulation of the graphical model selection problem is introduced. This allows for the development of a highly scalable uncertainty quantification method with theoretical safeguards. The fourth research aspect of the project examines the use of the methodology proposed in the previous three sub-components to an application in the area of climate change, where high dimensional covariance estimation is required. The proposal also has a significant teaching and outreach component which aims to introduce statistics to aspiring young scientists at various stages of their undergraduate and graduate studies.The availability of high-throughput data from various applications, including genomics, environmental sciences and others, has created an urgent need for methodology and tools for analyzing high dimensional data. Extracting and making sense of the many complex relationships and multivariate dependencies in the data and developing principled inferential procedures is one of the major challenges facing statisticians and data scientists. The theoretical and methodological work proposed in this project is motivated by applications and interdisciplinary collaborations in fields as diverse as the earth and environmental sciences, genomics and cancer research, and the social sciences. In genomics for instance, one is often interested to know how various genes are associated, and how these associations differ between an experimental (diseased) and control group. Gene regulatory networks also serve as important tools to study the evolutions of diseases. In the context of the climate change debate, modeling temperature at different points on the globe requires parsimonious modeling of the way in which these variables are related. Modeling correlations also arises naturally in material sciences and engineering where one is interested in seeing how different atomic particles interact when new materials are produced. Hence the proposed project for estimating correlations in very high dimensional settings will have widespread applications, since understanding associations/relationships between many variables is an endeavor that is common to many scientific disciplines. The proposed work, though firmly rooted in the statistical sciences, is very much interdisciplinary, and involves collaborations and partnerships between statisticians/data scientists and biomedical scientists, engineers and earth scientists.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Scalable methods for discovering multivariate dependencies in high dimensional data.
-
批准号:1916787
-
项目类别:Continuing Grant
-
资助金额:$29.32万
-
财政年份:2017
-
负责人:Balakanapathy Rajaratnam
-
依托单位:
Collaborative Research: Objective Bayesian Model Selection and Estimation in High Dimensional Statistical Models
-
批准号:1106642
-
项目类别:Standard Grant
-
资助金额:$9.92万
-
财政年份:2011
-
负责人:Balakanapathy Rajaratnam
-
依托单位:
CMG Collaborative Research: Efficient high dimensional Bayesian methods for climate field reconstruction
-
批准号:1025465
-
项目类别:Standard Grant
-
资助金额:$35.46万
-
财政年份:2010
-
负责人:Balakanapathy Rajaratnam
-
依托单位:
Collaborative Research: P2C2--Multiproxy Reconstructions as A Missing-Data Problem: New Techniques and their Application to Regional Climates of the Past Millennium
-
批准号:1003823
-
项目类别:Standard Grant
-
资助金额:$19.79万
-
财政年份:2010
-
负责人:Balakanapathy Rajaratnam
-
依托单位:
Exploring and detecting complex multivariate dependencies through sparse graphical models
-
批准号:0906392
-
项目类别:Standard Grant
-
资助金额:$10.38万
-
财政年份:2009
-
负责人:Balakanapathy Rajaratnam
-
依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: