课题基金 / 基金详情

Statistical Methods for Genomic Analysis of Species Divergences

Statistical Methods for Genomic Analysis of Species Divergences
物种差异基因组分析的统计方法
批准号:
BB/K000896/1
负责人:
Ziheng Yang
金额:
$42.57万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --

项目摘要

项目成果

Ziheng Yang的其他基金

相似基金

相关文献

中文摘要
翻译
我们的进化史写在我们的基因组里。通过比较不同物种的DNA序列,我们可以找出物种之间的关系。通过比较来自同一物种的多个个体的DNA序列,我们可以估计该物种的种群规模,并推断该物种的人口变化(如种群瓶颈)。这类研究属于遗传学和群体遗传学的范畴。来自几个密切相关物种的多个个体的基因组序列数据允许在遗传学和群体遗传学的界面上进行强有力的推断。人们可以利用这些数据来估计物种的分歧时间和祖先种群的大小,解释谱系排序,并检测物种形成时的基因流或测试不同的物种形成模型。这些数据还可以用来界定物种(例如,决定样本个体属于一个还是两个物种)。为了实现这些目标,强大的统计方法和计算算法是必要的。在这个项目中,我们将在两个完善的统计框架内实现这些方法:最大似然和贝叶斯推理。我们将开发最大似然方法来估计种群之间的迁移率,并设计似然比检验来检验物种形成时是否存在基因流(即物种形成是否干净)。我们将实施的模型,使迁移率随着时间的推移,因为物种的分歧。这些方法将有助于测试不同的物种形成模型,如异地和近地物种形成。计算困难将限制我们的可能性方法,在每个采样位点2或3个序列。然而,这些方法可以容纳大量的基因座(实际上是整个基因组),并且在某些基因座处具有种群数据,在其他基因座处具有物种数据,强大的推断是可行的。我们将使用计算机模拟来检查新方法的统计特性,并将这些方法应用于原始人的基因组数据集。我们将介绍对贝叶斯模型比较方法的重大改进和扩展,以使用基因组序列数据来划分物种。一年前发表的(Yang和Rannala 2010 Proc Natl Acad Sci USA 107:9264-9269),这种方法引起了进化生物学家的广泛关注。这使用了一种称为可逆跳马尔可夫链蒙特卡罗(reversible-jump Markov chain Monto Carlo,rjMCMC)的算法来对不同的物种划界模型进行采样,例如单物种模型(假设所有采样的个体都来自一个物种)和两个物种模型(假设采样的个体来自两个不同的物种)。然而,我们目前在计算机程序BPP中的实现具有严重的局限性,并且在中间或大型数据集中效率低下。该项目的一个主要目标是改进rjMCMC算法,使该程序成为可行的分析大型基因组规模的数据集。我们亦会将程式平行化,以提高计算效率。
英文摘要
Our evolutionary history is written in our genomes. By comparing DNA sequences from different species we can work out how the species are related. By comparing the DNA sequences of multiple individuals from the same species, we can estimate the population size and infer demographic changes (such as population bottleneck) of the species. Such studies fall into the domains of phylogenetics and population genetics. Genomic sequence data from multiple individuals of several closely related species allow powerful inference at the interface of phylogenetics and population genetics. One can use such data to estimate species divergence times and ancestral population sizes, accounting for lineage sorting, and to detect gene flow at the time speciation or to test different models of speciation Such data also allow delimitation of species (for example, to decide whether the sampled individuals belong to one or two species).To achieve those goals, powerful statistical methods and computational algorithms are necessary. In this project we will implement such methods within two well-established statistical frameworks: maximum likelihood and Bayesian inference. We will develop maximum likelihood methods for estimating migration rates between populations, and design likelihood ratio tests to test whether there is gene flow at the time of speciation (that is, whether speciation is clean). We will implement models that allow the migration rate to decrease over time since species divergence. Those methods will be useful for testing different speciation models such as allopatric and parapatric speciation. Computational difficulties will limit our likelihood methods to 2 or 3 sequences at each sampled locus. However the methods can accommodate a huge number of loci (indeed the whole genome), and with population data at some loci and species data at other loci, powerful inference is feasible. We will use computer simulations to examine the statistical properties of the new methods, and apply the methods to genomic datasets from the hominoids.We will introduce significant improvements and extensions to a Bayesian model-comparison approach to delimiting species using genomic sequence data. Published a year ago (Yang and Rannala 2010 Proc Natl Acad Sci USA 107:9264-9269), this method has attracted much attention among evolutionary biologists. This uses an algorithm called reversible-jump Markov chain Monto Carlo (rjMCMC) to sample different species-delimitation models, such as the one-species model (which assumes that all sampled individuals are from one single species) and the two species model (which assumes that the sampled individuals are from two distinct species). However, our current implementation in the computer program BPP has serious limitations and is inefficient in intermediate or large datasets. A major objective of this project is to improve the rjMCMC algorithm so that the program becomes feasible for analysis of large genomic-scale datasets. We will also parallelize the programs to improve the computational efficiency.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1111/2041-210x.12356
发表时间: 2015-06-01
期刊: METHODS IN ECOLOGY AND EVOLUTION
影响因子: 6.6
作者: [Liu, Junfeng, Zhang, De-Xing, Yang, Ziheng]
通讯作者: Yang, Ziheng
DOI: 10.1038/s41559-017-0280-x
发表时间: 2017-10
期刊: Nature ecology & evolution
影响因子: 16.8
作者: [Nascimento FF, Reis MD, Yang Z]
通讯作者: Yang Z
DOI: 10.1093/sysbio/syt049
发表时间: 2014-01-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者: [Leache, Adam D., Harris, Rebecca B., Yang, Ziheng]
通讯作者: Yang, Ziheng
DOI: 10.1093/sysbio/syu052
发表时间: 2014-11
期刊: Systematic biology
影响因子: 6.5
作者: [Zhang C, Rannala B, Yang Z]
通讯作者: Yang Z
共 8 条
    Efficient computational technologies to resolve the Timetree of Life: from ancient DNA to species-rich phylogenies
    • 批准号:
      BB/Y004132/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $64.4万
    • 财政年份:
      2024
    • 负责人:
      Ziheng Yang
    • 依托单位:
    PAML 5: A friendly and powerful bioinformatics resource for phylogenomics
    • 批准号:
      BB/X018571/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $63.77万
    • 财政年份:
      2024
    • 负责人:
      Ziheng Yang
    • 依托单位:
    NSFDEB-NERC: Integrating computational, phenotypic, and population-genomic approaches to reveal processes of cryptic speciation and gene flow in Madag
    • 批准号:
      NE/X002071/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $30.95万
    • 财政年份:
      2023
    • 负责人:
      Ziheng Yang
    • 依托单位:
    Bayesian inference of the mode of speciation and gene flow using genomic data
    • 批准号:
      BB/X007553/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $82.61万
    • 财政年份:
      2023
    • 负责人:
      Ziheng Yang
    • 依托单位:
    国内基金
    海外基金
    Computational Methods for Analyzing Toponome Data