Multivariate Statistical Methods for Genomic Data Integration
Multivariate Statistical Methods for Genomic Data Integration
批准号:
1457935
负责人:
Debashis Ghosh
金额:
$47.04万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-01 至 2017-05-31
中文摘要
该项目解决了许多使用基因组数据的数据分析师所面临的一个关键建模问题。对于一组个体或观察,生成了许多不同类型的高通量实验数据集,然后问题就变成了如何对这些数据建模。在许多问题中,目标是优先考虑希望研究基因组的哪些部分。虽然通常假设不同的数据类型在无条件或条件意义上是线性相关的,但在许多设置中,相关性的性质是未知的。本研究着重于高维基因组数据的多变量分析方法,放松线性假设。在这个项目的过程中,将研究两类问题。第一种是隐马尔可夫模型,第二种是多重测试程序,它们在基因组数据集上的应用已经很普遍。该项目提出了两种方法的新颖多元扩展,其目标是在大数据集上具有可靠的理论统计原理,同时在计算上可行。该方法将使用几个真实数据集以及通过模拟研究进行评估。这项工作将涉及统计学家和生物学家之间的相互作用。这项工作的更广泛用途将是在任何生物学环境中优先考虑后续研究的分子。这将有助于研究疾病过程的生物学家和科学家寻找新的治疗靶点或进一步推进基本的病因学理解。该项目的教育目标包括宾夕法尼亚州立大学研究生的新课程组成部分和统计学研究生的培训。
英文摘要
This project addresses a key modeling issue faced by many data analysts working with genomic data. For a set of individuals or observations, many different types of high throughput experimental datasets are generated, and the question then becomes how to model these data. In many problems, the goal is to prioritize which parts of the genome one wishes to study. While it is commonly assumed that the different data types are linearly correlated in either an unconditional or conditional sense, in many settings the nature of the correlation is unknown. This research focuses on multivariate methods of analysis with high-dimensional genomic data that relax the linearity assumption. Two classes of problems will be studied during the course of the project. The first is Hidden Markov Models and the second is multiple testing procedures, whose use have become commonplace with genomic datasets. This project proposes novel multivariate extensions of both types of method with a goal of being characterized by sound theoretical statistical principles while simultaneously being computationally feasible on big datasets. The methodology will be evaluated using several real datasets as well as through simulation studies.This work will involve an interplay between statisticians and biologists. The broader use of this work will be to prioritize molecules for follow-up studies in any biological setting. It will be useful for biologists and scientists studying disease processes who wish to find new therapeutic targets or further advance basic etiological understanding. The educational goals of the project include new course components for graduate students at Penn State and training of graduate students in Statistics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Empirical and Causal Models for Heterogeneous Data Fusion
-
批准号:2149492
-
项目类别:Standard Grant
-
资助金额:$28.15万
-
财政年份:2022
-
负责人:Debashis Ghosh
-
依托单位:
New Methods in High-Dimensional Causal Inference
-
批准号:1914937
-
项目类别:Standard Grant
-
资助金额:$14.98万
-
财政年份:2019
-
负责人:Debashis Ghosh
-
依托单位:
Multivariate Statistical Methods for Genomic Data Integration
-
批准号:1262538
-
项目类别:Continuing Grant
-
资助金额:$54.56万
-
财政年份:2013
-
负责人:Debashis Ghosh
-
依托单位:
海外基金