Multivariate Statistical Methods for Genomic Data Integration
Multivariate Statistical Methods for Genomic Data Integration
批准号:
1457935
负责人:
Debashis Ghosh
金额:
$47.04万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-01 至 2017-05-31
中文摘要
该项目解决了许多从事基因组数据工作的数据分析师所面临的关键建模问题。对于一组个人或观察,会生成许多不同类型的高通量实验数据集,然后问题就变成了如何对这些数据建模。在许多问题中,目标是确定人们希望研究基因组的哪些部分的优先顺序。虽然通常假设不同的数据类型在无条件或条件意义上是线性相关的,但在许多设置中,相关性的性质是未知的。这项研究集中在用高维基因组数据进行分析的多变量方法,放松了线性假设。在项目过程中,将研究两类问题。第一个是隐马尔可夫模型,第二个是多重测试程序,它们在基因组数据集中的使用已经变得司空见惯。该项目提出了两种方法的新的多变量扩展,目标是以合理的理论统计原理为特征,同时在大数据集上计算可行。该方法将使用几个真实的数据集以及通过模拟研究进行评估。这项工作将涉及统计学家和生物学家之间的相互作用。这项工作的更广泛用途将是为任何生物学环境中的后续研究确定分子的优先顺序。对于研究疾病过程的生物学家和科学家来说,这将是有用的,他们希望找到新的治疗靶点或进一步推进基本的病因学理解。该项目的教育目标包括为宾夕法尼亚州立大学的研究生提供新的课程内容和对研究生进行统计学培训。
英文摘要
This project addresses a key modeling issue faced by many data analysts working with genomic data. For a set of individuals or observations, many different types of high throughput experimental datasets are generated, and the question then becomes how to model these data. In many problems, the goal is to prioritize which parts of the genome one wishes to study. While it is commonly assumed that the different data types are linearly correlated in either an unconditional or conditional sense, in many settings the nature of the correlation is unknown. This research focuses on multivariate methods of analysis with high-dimensional genomic data that relax the linearity assumption. Two classes of problems will be studied during the course of the project. The first is Hidden Markov Models and the second is multiple testing procedures, whose use have become commonplace with genomic datasets. This project proposes novel multivariate extensions of both types of method with a goal of being characterized by sound theoretical statistical principles while simultaneously being computationally feasible on big datasets. The methodology will be evaluated using several real datasets as well as through simulation studies.This work will involve an interplay between statisticians and biologists. The broader use of this work will be to prioritize molecules for follow-up studies in any biological setting. It will be useful for biologists and scientists studying disease processes who wish to find new therapeutic targets or further advance basic etiological understanding. The educational goals of the project include new course components for graduate students at Penn State and training of graduate students in Statistics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Empirical and Causal Models for Heterogeneous Data Fusion
-
批准号:2149492
-
项目类别:Standard Grant
-
资助金额:$28.15万
-
财政年份:2022
-
负责人:Debashis Ghosh
-
依托单位:
New Methods in High-Dimensional Causal Inference
-
批准号:1914937
-
项目类别:Standard Grant
-
资助金额:$14.98万
-
财政年份:2019
-
负责人:Debashis Ghosh
-
依托单位:
Multivariate Statistical Methods for Genomic Data Integration
-
批准号:1262538
-
项目类别:Continuing Grant
-
资助金额:$54.56万
-
财政年份:2013
-
负责人:Debashis Ghosh
-
依托单位:
海外基金