课题基金 / 基金详情

项目摘要

项目成果

Yun S Song的其他基金

相似基金

相关文献

中文摘要
翻译
 描述(由申请人提供): 在过去的几年里,DNA测序技术的进步极大地增加了基因组变异数据的可用性。这一发展为了解人类生物学和疾病风险的遗传基础提供了一个强有力的窗口。为了促进实现这一目标,至关重要的是开发有效的分析方法,使研究人员能够更充分地利用基因组数据中的信息,并考虑比以前更复杂的模型。该项目的中心目标是通过执行以下具体目标来应对这一重要挑战:在目标1中,我们将通过扩展我们正在进行的关于合并隐马尔可夫模型的工作,并将其应用于大规模数据,来开发用于全基因组群体基因组分析的有效推理工具。我们开发的方法将使研究人员能够在一般人口统计模型下分析大样本,该模型涉及多个人口,人口分裂,迁移和混合,以及可变的有效人口规模和时间样本(古代DNA)。在大多数复杂的群体遗传模型中,多位点全似然计算往往是不可行的。为了解决这个问题,我们将在Aim 2中开发一种新的无可能性推理框架,用于通过应用高度活跃的机器学习研究领域(称为深度学习)进行群体基因组分析。我们将应用该方法在人口基因组学的各种参数估计和分类问题,特别是联合推理的选择和人口。除了进行技术研究外,我们还将开发一个有用的软件包,使人口基因组学社区的研究人员能够在自己的研究中利用深度学习。利用全基因组规模的时间序列遗传变异数据来推断等位基因频率随时间的变化变得越来越流行。这一发展创造了新的机会,以确定选择压力下的基因组区域,并估计其相关的健身参数。在目标3中,我们将开发新的统计方法,以充分利用这种新的数据源在短期和长期的进化时间尺度。具体来说,我们将开发和应用有效的统计推断方法,用于分析来自实验进化和古代DNA样本的时间序列基因组变异数据。将为每个具体目标开发有用的开放源码软件。该项目开发的新方法将有助于在全基因组规模上分析和解释遗传变异数据。
英文摘要
 DESCRIPTION (provided by applicant): Technological advances in DNA sequencing have dramatically increased the availability of genomic variation data over the past few years. This development offers a powerful window into understanding the genetic basis of human biology and disease risk. To facilitate achieving this goal, it is crucial to develop efficient analytical methods that will allow researchers to more fuly utilize the information in genomic data and consider more complex models than previously possible. The central goal of this project is to tackle this important challenge, by carrying out te following Specific Aims: In Aim 1, we will develop efficient inference tools for whole-genome population genomic analysis by extending our ongoing work on coalescent hidden Markov models and apply them to large-scale data. The methods we develop will enable researchers to analyze large samples under general demographic models involving multiple populations with population splits, migration, and admixture, as well as variable effective population sizes and temporal samples (ancient DNA). Multi-locus full-likelihood computation is often prohibitive in most population genetic models with high complexity. To address this problem, we will develop in Aim 2 a novel likelihood-free inference framework for population genomic analysis by applying a highly active area of machine learning research called deep learning. We will apply the method to various parameter estimation and classification problems in population genomics, particularly joint inference of selection and demography. In addition to carrying out technical research, we will develop a useful software package that will allow researchers from the population genomics community to utilize deep learning in their own research. It is becoming increasingly more popular to utilize time-series genetic variation data at the whole-genome scale to infer allele frequency changes over a time course. This development creates new opportunities to identify genomic regions under selective pressure and to estimate their associated fitness parameters. In Aim 3, we will develop new statistical methods to take full advantage of this novel data source at both short and long evolutionary timescales. Specifically, we will develop and apply efficient statistical inference methods for analyzing time-series genomic variation data from experimental evolution and ancient DNA samples. Useful open-source software will be developed for each specific aim. The novel methods developed in this project will help to analyze and interpret genetic variation data at the whole-genome scale.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Robust and efficient statistical inference methods for genomics
Robust and efficient statistical inference methods for genomics
Robust and efficient statistical inference methods for genomics
Robust and efficient statistical inference methods for genomics
海外基金