What can we learn about the distribution of fitness effects of new mutations from DNA sequence data?

What can we learn about the distribution of fitness effects of new mutations from DNA sequence data?
复制标题

DOI:
10.1098/rstb.2009.0266
复制
发表时间:
2010-04-27
影响因子:
6.3
通讯作者:
Eyre-Walker, Adam
Eyre-Walker, Adam
中科院分区:
生物学1区
文献类型:
--
作者:
Keightley, Peter D.;Eyre-Walker, Adam

文献摘要

被引文献

相似文献

我们研究了几个问题,涉及到从群体样本中核苷酸频率的分布推断新突变的适合度效应分布(DFE)。如果有固定的测序工作,我们发现最好的策略是对适度数量的等位基因进行测序(约.10)。如果有完整的基因组信息,参数估计的准确性会随着测序的等位基因数量的增加而增加,但回报会减少。在人类和果蝇等生物体中,不太可能可靠地估计单个基因的DFE,除非基因非常大,我们对数百甚至数千个等位基因进行了测序。我们考虑包含几个离散的突变类别的模型,其中分配给每个类别的选择强度和密度可以变化。除非对数百个等位基因进行测序,否则具有三个类别的模型与四个类别的模型几乎一样符合。为了准确估计分布的均值和方差,需要对大量的等位基因进行测序。因此,估计复杂的DFE可能很困难。最后,我们检查涉及略微有利的突变的模型。我们证明,如果假设突变是无条件有害的,那么选择的绝对强度分布是可以很好地估计的。
We investigate several questions concerning the inference of the distribution of fitness effects (DFE) of new mutations from the distribution of nucleotide frequencies in a population sample. If a fixed sequencing effort is available, we find that the optimum strategy is to sequence a modest number of alleles (approx. 10). If full genome information is available, the accuracy of parameter estimates increases as the number of alleles sequenced increases, but with diminishing returns. It is unlikely that the DFE for single genes can be reliably estimated in organisms such as humans and Drosophila, unless genes are very large and we sequence hundreds or perhaps thousands of alleles. We consider models involving several discrete classes of mutations in which the selection strength and density apportioned to each class can vary. Models with three classes fit almost as well as four class models unless many hundreds of alleles are sequenced. Large numbers of alleles need to be sequenced to accurately estimate the distribution's mean and variance. Estimating complex DFEs may therefore be difficult. Finally, we examine models involving slightly advantageous mutations. We show that the distribution of the absolute strength of selection is well estimated if mutations are assumed to be unconditionally deleterious.