A General Method for Calculating Likelihoods Under the Coalescent Process

A General Method for Calculating Likelihoods Under the Coalescent Process
复制标题

DOI:
10.1534/genetics.111.129569
复制
发表时间:
2011-11-01
期刊:
影响因子:
3.3
通讯作者:
Barton, N. H.
Barton, N. H.
中科院分区:
生物学2区
文献类型:
--
作者:
Lohse, K.;Harrison, R. J.;Barton, N. H.

文献摘要

被引文献

相似文献

基因组数据的分析需要一种有效的方法来计算非常大量的基因座的可能性。我们描述了一种寻找家系分布的一般方法:我们允许Deme之间的迁移,Deme的分裂[如在隔离-迁移(IM)模型中],以及连锁的基因座之间的重组。这些过程由分支长度的母函数的一组线性递归来描述。在无限位置模型下,通过对这个母函数的求导,可以求出突变的任何构型的概率。这样的计算对于小样本基因组是可行的:作为一个例子,我们展示了如何在两维IM模型下显式地推导出三个基因的发生函数。这一推导是使用MATHEMICAL自动完成的。给定来自大量未连接和未重组的序列块的数据,这些结果可用于通过将所有相关突变构型的概率制表,然后在多个基因座上相乘来找到模型参数的最大似然估计。通过将该方法应用于模拟数据和Wang和Hey(2010)之前分析的包含26,141个从果蝇和黑腹果蝇中采样的基因座的数据集,证明了该方法的可行性。我们的结果表明,只要每个序列块的采样个体和突变的数量很小,这种似然计算就可以扩展到基因组数据。
Analysis of genomic data requires an efficient way to calculate likelihoods across very large numbers of loci. We describe a general method for finding the distribution of genealogies: we allow migration between demes, splitting of demes [as in the isolation-with-migration (IM) model], and recombination between linked loci. These processes are described by a set of linear recursions for the generating function of branch lengths. Under the infinite-sites model, the probability of any configuration of mutations can be found by differentiating this generating function. Such calculations are feasible for small numbers of sampled genomes: as an example, we show how the generating function can be derived explicitly for three genes under the two-deme IM model. This derivation is done automatically, using Mathematica. Given data from a large number of unlinked and nonrecombining blocks of sequence, these results can be used to find maximum-likelihood estimates of model parameters by tabulating the probabilities of all relevant mutational configurations and then multiplying across loci. The feasibility of the method is demonstrated by applying it to simulated data and to a data set previously analyzed by Wang and Hey (2010) consisting of 26,141 loci sampled from Drosophila simulans and D. melanogaster. Our results suggest that such likelihood calculations are scalable to genomic data as long as the numbers of sampled individuals and mutations per sequence block are small.