ESTIMATING EFFECTIVE POPULATION-SIZE FROM SAMPLES OF SEQUENCES - A BOOTSTRAP MONTE-CARLO INTEGRATION METHOD

ESTIMATING EFFECTIVE POPULATION-SIZE FROM SAMPLES OF SEQUENCES - A BOOTSTRAP MONTE-CARLO INTEGRATION METHOD
复制标题

DOI:
10.1017/s0016672300030962
复制
发表时间:
1992-12-01
期刊:
影响因子:
1.5
通讯作者:
FELSENSTEIN, J
FELSENSTEIN, J
中科院分区:
生物学4区
文献类型:
--
作者:
FELSENSTEIN, J

文献摘要

被引文献

相似文献

我们希望使用最大似然估计参数,如有效群体大小N(e),或者,如果我们不知道突变率,则每个位点的突变率与有效群体大小的乘积4N(e)μ。为了计算取自随机交配群体的未重组核苷酸序列样本的可能性,有必要对可能产生序列的所有系谱进行求和,计算每一个系谱产生序列的概率,并根据其先验概率对每一个系谱进行加权。系谱在树拓扑和分支长度上不同。虽然可能性和先验是直接计算的,但对所有系谱求和乍一看似乎是无可救药的困难。本文报道,它是可能的,进行蒙特卡罗积分近似估计的可能性。该方法使用自举抽样的网站,以创建数据集,其中每个最大似然树估计。假设所得到的树是从一个分布中采样的,该分布的高度与完整数据的似然曲面成比例。这取决于一个尚未证明的定理,但如果序列不短,这个定理似乎是正确的。可以使用所得的估计似然曲线来对感兴趣的参数N(e)或4N(e)μ进行最大似然估计。该方法需要至少100倍的计算工作量所需的最大似然估计的一个遗传学,但在今天的工作站是实用的。该方法目前没有任何处理重组的方法。
We would like to use maximum likelihood to estimate parameters such as the effective population size N(e), or, if we do not know mutation rates, the product 4N(e)mu of mutation rate per site and effective population size. To compute the likelihood for a sample of unrecombined nucleotide sequences taken from a random-mating population it is necessary to sum over all genealogies that could have led to the sequences, computing for each one the probability that it would have yielded the sequences, and weighting each one by its prior probability. The genealogies vary in tree topology and in branch lengths. Although the likelihood and the prior are straightforward to compute, the summation over all genealogies seems at first sight hopelessly difficult. This paper reports that it is possible to carry out a Monte Carlo integration to evaluate the likelihoods approximately. The method uses bootstrap sampling of sites to create data sets for each of which a maximum likelihood tree is estimated. The resulting trees are assumed to be sampled from a distribution whose height is proportional to the likelihood surface for the full data. That it will be so is dependent on a theorem which is not proven, but seems likely to be true if the sequences are not short. One can use the resulting estimated likelihood curve to make a maximum likelihood estimate of the parameter of interest, N(e) or of 4N(e)mu. The method requires at least 100 times the computational effort required for estimation of a phylogeny by maximum likelihood, but is practical on today's work stations. The method does not at present have any way of dealing with recombination.