Bayes empirical Bayes inference of amino acid sites under positive selection

Bayes empirical Bayes inference of amino acid sites under positive selection
复制标题

DOI:
10.1093/molbev/msi097
复制
发表时间:
2005-04-01
影响因子:
10.7
通讯作者:
Nielsen, R
Nielsen, R
中科院分区:
生物学1区
文献类型:
--
作者:
Yang, ZH;Wong, WSW;Nielsen, R

文献摘要

被引文献

相似文献

基于密码子的替换模型已被广泛用于蛋白质编码DNA序列比较分析中确定正选择下的氨基酸位点。非同义-同义取代率比(d(N)/d(S),表示为ω)用作蛋白质水平上的选择压力的量度,ω> 1表示正选择。统计分布用于模拟位点之间的ω的变化,允许位点的子集具有ω> 1,而序列的其余部分可以处于ω < 1的纯化选择下。然后使用经验贝叶斯(EB)方法来计算站点来自ω> 1的站点类的后验概率。然而,目前的实现,使用朴素的EB(NEB)的方法,并没有考虑到模型参数的最大似然估计的采样误差,如比例和Ω比的网站类。在缺乏信息的小数据集,这种方法可能会导致不可靠的后验概率计算。在本文中,我们开发了一个贝叶斯经验贝叶斯(BEB)的方法来解决这个问题,它分配一个先验模型参数,并整合了他们的不确定性。在真实的数据集和模拟数据集上对新方法和旧方法进行了比较。结果表明,在小的数据集,新的BEB方法不会产生假阳性的旧NEB方法,而在大的数据集,它保留了良好的权力NEB方法推断积极选择的网站。
Codon-based substitution models have been widely used to identify amino acid sites under positive selection in comparative analysis of protein-coding DNA sequences. The non synonymous-synonymous substitution rate ratio (d(N)/d(S), denoted omega) is used as a measure of selective pressure at the protein level, with omega > 1 indicating positive selection. Statistical distributions are used to model the variation in omega among sites, allowing a subset of sites to have omega > 1 while the rest of the sequence may be under purifying selection with omega < 1. An empirical Bayes (EB) approach is then used to calculate posterior probabilities that a site comes from the site class with omega > 1. Current implementations, however, use the naive EB (NEB) approach and fail to account for sampling errors in maximum likelihood estimates of model parameters, such as the proportions and omega ratios for the site classes. In small data sets lacking information, this approach may lead to unreliable posterior probability calculations. In this paper, we develop a Bayes empirical Bayes (BEB) approach to the problem, which assigns a prior to the model parameters and integrates over their uncertainties. We compare the new and old methods on real and simulated data sets. The results suggest that in small data sets the new BEB method does not generate false positives as did the old NEB approach, while in large data sets it retains the good power of the NEB approach for inferring positively selected sites.