III: Medium: Detecting Low Dimensional Structures in Genomic Data
III: Medium: Detecting Low Dimensional Structures in Genomic Data
批准号:
1705197
负责人:
Eleazar Eskin
金额:
$119.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-15 至 2022-07-31
中文摘要
新的测序技术使基因组学成为一门大数据科学。这些数据具有复杂性,代表了许多变量。在试图从基因组序列中获取生物信息时,往往需要降低复杂性。有许多不同的计算方法可供使用,但由于对数据所做的假设,这些方法通常会引入错误。该项目将导致针对所收集的基因组数据类型开发新的方法。其中一种类型的数据代表DNA序列,另一种来自基因表达时对序列的自然修改。这些新方法将通过在统计框架中正确模拟这些数据的独特属性,更准确地确定这两种数据类型中的重要差异。在该项目期间开发的方法将对基因组学领域产生重大影响,研究人员可能会在那里发现复杂疾病的遗传基础。该项目的更广泛影响是更深入地了解复杂疾病的遗传基础,通过公共网络服务器和用于学术研究和教育目的的软件工具传播新方法,并培训本科生、研究生和博士后学者。特别是,该项目将通过夏季强化计划为代表不足的群体提供培训,招募传统上在STEM领域代表不足的少数群体。从高维基因组数据中发现低维结构在基因组研究中是一个非常重要的过程,因为这种结构可能推断基因组数据中未知的混杂因素以及其他重要的数据属性,如个体的种族。基因组学中普遍使用的降维方法有几种,它们可能不能从基因组数据中生成准确的低维结构,因为它们对统计模型的基本假设经常在数据中被违反。该项目建议开发针对基因组数据的降维方法,特别是针对甲基化和基因数据。这些方法将结合基因组数据中存在的独特性质,如基因数据的离散性质和相关结构,以及不同细胞类型和组织中不同的甲基化模式。本项目还将使用随机矩阵理论分析新方法的渐近行为。将使用三种策略来验证这些方法。首先,对于所有基因组学应用,都有数据集,其中有黄金标准信息;其次,将使用基于基因组学社区当前实践的模拟数据来执行评估基因组学应用。例如,在社区中,通过结合参考数据集(如1000基因组计划)中已知祖先的个体的基因类型来模拟混合个体的遗传学是标准的。第三,该团队将通过使用各种生成模型生成模拟数据来评估一般算法,以验证算法具有预期的渐近行为,并检查这些算法在违反其假设时的表现。这些方法将通过改进现有的低维方法对统计领域做出贡献,并通过发布软件工具对基因组学领域做出贡献。该项目的更广泛影响是更深入地了解复杂疾病的遗传基础,通过公共网络服务器和用于学术研究和教育目的的软件工具分发方法,并培训本科生、研究生和博士后学者。特别是,该项目将通过夏季强化方案向任职人数不足的群体提供培训,招募传统上在STEM领域任职人数不足的少数民族。
英文摘要
New sequencing technologies have made genomics a big data science. These data have complexity and represent many variables. In trying to get biological information from genomic sequence, it is often necessary to reduce the complexity. There are a number of different approaches to use computationally, but these often introduce errors because of assumptions made about the data. This project will lead to the development of novel approaches specific to the type of genomic data collected. One of these types of data represents the DNA sequence and the other comes from natural modifications to the sequence when genes are expressed. These new methods will identify important differences more accurately in the two data types by correctly modeling unique properties of these data in a statistical framework. Methods developed during this project will have a great impact on the genomics field, where researchers may discover the genetic basis of complex diseases. The broader impacts of this project are gaining a deeper insight into the genetic basis of complex diseases, distributing the novel methods through public webservers and software tools for academic research and educational purposes, and training undergraduate students, graduate students, and postdoctoral scholars. In particular, this project will provide training to underrepresented groups with a summer intensive program that recruits minorities traditionally underrepresented in STEM fields.Discovering a low dimensional structure from the high dimensional genomic data is a very important procedure in genomic studies because this structure may infer unknown confounding factors in genomic data as well as other important properties of data such as ethnicity of individuals. There are several dimensionality reduction methods prevalently used in the genomics, they may not generate an accurate low dimensional structure from genomic data because their underlying assumption on the statistical model is often violated in the data. This project proposes to develop dimensionality reduction methods aimed for genomic data, especially for methylation and genotype data. These methods will incorporate unique properties present in genomic data such as the discrete nature and correlation structure of genotype data, and different methylation patterns across different cell types and tissues. This project will also analyze asymptotic behavior of the novel methods using random matrix theory. Three strategies will be used to validate the methods. First, for all genomics applications, there are datasets where there is gold standard information, Second, simulated data based on current practices in the genomics community will be used to perform evaluate genomics applications. For example, it is standard in the community to simulate the genetics of admixed individuals by combining the genotypes of individuals of known ancestry from a reference dataset such as the 1000 Genomes project. Third, the team will evaluate the general algorithms by generating simulated data using various generative models to validate that the algorithms have the asymptotic behavior expected and also examine how these algorithms perform when their assumptions are violated. The methods will contribute both to the statistical field by improving current low dimensionality methods and to the genomics field by releasing software tools. The broader impacts of this project are gaining a deeper insight into the genetic basis of complex diseases, distributing the methods through public webservers and software tools for academic research and educational purposes, and training undergraduate students, graduate students, and postdoctoral scholars. In particular, this project will provide training to underrepresented groups with a summer intensive program that recruits minorities traditionally underrepresented in STEM fields.
期刊论文(28)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1038/s41598-020-67513-5
发表时间:
2020-07-03
期刊:
SCIENTIFIC REPORTS
影响因子:
4.6
作者:
[Alvarez, Marcus, Rahmani, Elior, Pajukanta, Paivi]
通讯作者:
Pajukanta, Paivi
DOI:
10.1371/journal.pgen.1008481
发表时间:
2019-12-01
期刊:
PLOS GENETICS
影响因子:
4.5
作者:
[Zou, Jennifer, Hormozdiari, Farhad, Eskin, Eleazar]
通讯作者:
Eskin, Eleazar
Contribution of common and rare variants to bipolar disorder susceptibility in extended pedigrees from population isolates.
人群分离株的扩展谱系中常见和罕见变异对双相情感障碍易感性的贡献。
DOI:
10.1038/s41398-020-0758-1
发表时间:
2020
期刊:
Translational psychiatry
影响因子:
6.8
作者:
[Sul,JaeHoon, Service,SusanK, Huang,AldenY, Ramensky,Vasily, Hwang,Sun-Goo, Teshiba,TerriM, Park,YoungJun, Ori,AnilPS, Zhang,Zhongyang, Mullins,Niamh, OldeLoohuis,LoesM, Fears,ScottC, Araya,Carmen, Araya,Xinia, Spesny,Mitzi, Bejaran]
通讯作者:
Bejaran
DOI:
10.1038/s41467-020-15652-8
发表时间:
2020-04-20
期刊:
NATURE COMMUNICATIONS
影响因子:
16.6
作者:
[Furman, Ori, Shenhav, Liat, Mizrahi, Itzhak]
通讯作者:
Mizrahi, Itzhak
DOI:
10.1371/journal.pcbi.1007556
发表时间:
2019-12-01
期刊:
PLOS COMPUTATIONAL BIOLOGY
影响因子:
4.3
作者:
[Li, Jiajin, Jew, Brandon, Sul, Jae Hoon]
通讯作者:
Sul, Jae Hoon
共 12 条
III: Medium: Causal inference in biobanks: Leveraging genetics to infer causal relationships using electronic health records
-
批准号:2106908
-
项目类别:Continuing Grant
-
资助金额:$119.99万
-
财政年份:2021
-
负责人:Eleazar Eskin
-
依托单位:
III:Small: Replication Studies for High Dimensional Data: Insights into Confounding and Heterogeneity
-
批准号:1910885
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2019
-
负责人:Eleazar Eskin
-
依托单位:
III: Small: Causal and Statistical Inference in the Presence of Confounding Factors
-
批准号:1320589
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2013
-
负责人:Eleazar Eskin
-
依托单位:
BSF:2012304:Methods for Preprocessing Population Sequence Data
-
批准号:1331176
-
项目类别:Standard Grant
-
资助金额:$4.0万
-
财政年份:2013
-
负责人:Eleazar Eskin
-
依托单位:
III: Medium: Meta-analysis reinterpreted using causal graphs
-
批准号:1302448
-
项目类别:Continuing Grant
-
资助金额:$112.08万
-
财政年份:2013
-
负责人:Eleazar Eskin
-
依托单位:
III: Medium: Private Identification of Relatives and Private GWAS: First Steps in the New Field of CryptoGenomics
-
批准号:1065276
-
项目类别:Standard Grant
-
资助金额:$70.0万
-
财政年份:2011
-
负责人:Eleazar Eskin
-
依托单位:
III: Small: Inference of Causal Regulatory Relationships from Genetic Studies
-
批准号:0916676
-
项目类别:Continuing Grant
-
资助金额:$49.94万
-
财政年份:2009
-
负责人:Eleazar Eskin
-
依托单位:
Collaborative Research: Design and Analysis of Compressed Sensing DNA Microarrays
-
批准号:0729049
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:2007
-
负责人:Eleazar Eskin
-
依托单位:
Collaborative Research: SEIII: Estimating Haplotype Frequencies
-
批准号:0731455
-
项目类别:Standard Grant
-
资助金额:$13.59万
-
财政年份:2007
-
负责人:Eleazar Eskin
-
依托单位:
Collaborative Research: SEIII: Estimating Haplotype Frequencies
-
批准号:0513612
-
项目类别:Standard Grant
-
资助金额:$29.5万
-
财政年份:2005
-
负责人:Eleazar Eskin
-
依托单位:
海外基金