Accuracy of coalescent likelihood estimates: Do we need more sites, more sequences, or more loci?

Accuracy of coalescent likelihood estimates: Do we need more sites, more sequences, or more loci?
复制标题

DOI:
10.1093/molbev/msj079
复制
发表时间:
2006-03-01
影响因子:
10.7
通讯作者:
Felsenstein, J
Felsenstein, J
中科院分区:
生物学1区
文献类型:
--
作者:
Felsenstein, J

文献摘要

被引文献

相似文献

本文用计算机模拟研究了有限大小的单个孤立总体的样本对θ = 4N(e)μ的估计精度。傅和李使用乐观的假设开发的公式可以很好地预测精度。他们的公式在精确度方面进行了重述,这里定义为变异系数平方的倒数。当抽样实体提供独立信息时,这应与抽样规模成比例。使用这些公式的准确性,估计θ的抽样策略可以进行调查。使用了两种成本模式,即每次基成本模式和每次读取成本模式。前者会导致我们倾向于拥有大量的基因座,每个基因座都有一个碱基长。后者,这是更现实的,使我们更喜欢有一个读每个位点和一个最佳的样本量,随着采样生物体的成本增加而下降。对于实际值,最佳样本量为8个或更少个体。这与Pluzhnikov和Donnelly对每碱基成本模型所获得的结果非常接近,评估了Theta的其他估计值。可以理解的是,考虑到收集更大样本所花费的资源使我们无法考虑更多的位点。沃特森的估计θ的效率也进行了检查,它被认为是合理有效的,当在整个人口的序列中每一代的突变体的数量小于2.5。
A computer simulation study has been made of the accuracy of estimates of Theta = 4N(e)mu from a sample from a single isolated population of finite size. The accuracies turn out to be well predicted by a formula developed by Fu and Li, who used optimistic assumptions. Their formulas are restated in terms of accuracy, defined here as the reciprocal of the squared coefficient of variation. This should be proportional to sample size when the entities sampled provide independent information. Using these formulas for accuracy, the sampling strategy for estimation of Theta can be investigated. Two models for cost have been used, a cost-per-base model and a cost-per-read model. The former would lead us to prefer to have a very large number of loci, each one base long. The latter, which is more realistic, causes us to prefer to have one read per locus and an optimum sample size which declines as costs of sampling organisms increase. For realistic values, the optimum sample size is 8 or fewer individuals. This is quite close to the results obtained by Pluzhnikov and Donnelly for a cost-per-base model, evaluating other estimators of Theta It can be understood by considering that the resources spent collecting larger samples prevent us from considering more loci. An examination of the efficiency of Watterson's estimator of Theta was also made, and it was found to be reasonably efficient when the number of mutants per generation in the sequence in the whole population is less than 2.5.