Not so different after all: A comparison of methods for detecting amino acid sites under selection

Not so different after all: A comparison of methods for detecting amino acid sites under selection
复制标题

DOI:
10.1093/molbev/msi105
复制
发表时间:
2005-05-01
影响因子:
10.7
通讯作者:
Frost, SDW
Frost, SDW
中科院分区:
生物学1区
文献类型:
--
作者:
Pond, SLK;Frost, SDW

文献摘要

被引文献

相似文献

我们考虑了三种方法来估计序列比对中每个位点的非同义和同义变化率,以确定正选择或负选择下的位点:(1)一套快速的基于似然的“计数方法”,采用单一的最可能的祖先重建,加权所有可能的祖先重建,或从祖先重建中取样;(2)随机效应似然(REL)方法,其根据预定义的分布对不同地点的非同义和同义比率的变化进行建模,其中使用经验贝叶斯方法推断单个地点处的选择压力;以及(3)固定效应似然(FEL)方法,其直接估计每个位点的非同义和同义替换率。所有这三种方法都采用了灵活的模型,核苷酸取代的偏见和变化,在两个非同义和同义取代率跨网站,促进方法之间的比较。我们证明,使用这些方法得到的结果显示广泛的协议,在I型和11型错误的水平和替代率的估计。计数方法非常适合于大型比对,对于大型比对,有很高的能力来检测阳性和阴性选择,但似乎低估了替代率。一个REL的方法,这是更计算密集型比计数方法,具有更高的功率比计数方法来检测选择的数据集的中间大小,但可能遭受较高的误报率为小数据集。一个自由电子激光的方法似乎捕捉率变化的模式比计数方法或随机效应模型,不遭受许多假阳性的随机效应模型的数据集,包括几个序列,并可以有效地并行化。我们的研究结果表明,以前报道的计数方法和随机效应模型获得的结果之间的差异,由于计数为基础的方法,目前的随机效应模型,以允许同义替换率的变化,并天真的应用随机效应模型的保守性的组合非常稀疏的数据集。我们证明了我们的方法从人类免疫缺陷病毒1型env和pol基因和模拟比对的序列数据。
We consider three approaches for estimating the rates of nonsynonymous and synonymous changes at each site in a sequence alignment in order to identify sites under positive or negative selection: (1) a suite of fast likelihood-based "counting methods" that employ either a single most likely ancestral reconstruction, weighting across all possible ancestral reconstructions, or sampling from ancestral reconstructions; (2) a random effects likelihood (REL) approach, which models variation in nonsynonymous and synonymous rates across sites according to a predefined distribution, with the selection pressure at an individual site inferred using an empirical Bayes approach; and (3) a fixed effects likelihood (FEL) method that directly estimates nonsynonymous and synonymous substitution rates at each site. All three methods incorporate flexible models of nucleotide substitution bias and variation in both nonsynonymous and synonymous substitution rates across sites, facilitating the comparison between the methods. We demonstrate that the results obtained using these approaches show broad agreement in levels of Type I and Type 11 error and in estimates of substitution rates. Counting methods are well suited for large alignments, for which there is high power to detect positive and negative selection, but appear to underestimate the substitution rate. A REL approach, which is more computationally intensive than counting methods, has higher power than counting methods to detect selection in data sets of intermediate size but may suffer from higher rates of false positives for small data sets. A FEL approach appears to capture the pattern of rate variation better than counting methods or random effects models, does not suffer from as many false positives as random effects models for data sets comprising few sequences, and can be efficiently parallelized. Our results suggest that previously reported differences between results obtained by counting methods and random effects models arise due to a combination of the conservative nature of counting-based methods, the failure of current random effects models to allow for variation in synonymous substitution rates, and the naive application of random effects models to extremely sparse data sets. We demonstrate our methods on sequence data from the human immunodeficiency virus type 1 env and pol genes and simulated alignments.