Accuracy and power of Bayes prediction of amino acid sites under positive selection

Accuracy and power of Bayes prediction of amino acid sites under positive selection
复制标题

DOI:
10.1093/oxfordjournals.molbev.a004152
复制
发表时间:
2002-06-01
影响因子:
10.7
通讯作者:
Yang, ZH
Yang, ZH
中科院分区:
生物学1区
文献类型:
--
作者:
Anisimova, M;Bielawski, JP;Yang, ZH

文献摘要

被引文献

相似文献

贝叶斯预测通过分配后验概率来量化不确定性。它被用来鉴定在反复多样性选择下的蛋白质中的氨基酸,所述反复多样性选择通过非同义(d(N))比同义(d(S))替换率更高或ω = d(N)/d(S)> 1来指示。参数估计的密码子取代模型下,假设几类不同的w比的网站的最大似然。贝叶斯定理被用来计算后验概率的每个网站属于这些网站类。这里.我们通过计算机模拟来评估正选择下氨基酸的贝叶斯预测的性能。我们通过预测的真正处于选择下的位点的比例来测量准确性,并通过该方法预测的真正积极选择的位点的比例来测量功效。较长序列的准确性略好,而功率在很大程度上不受序列长度增加的影响。中等或高度发散序列的准确度和功效均高于相似序列。我们发现,当数据只包含少数高度相似的序列时,准确度和功效低得令人无法接受。然而,对大量谱系进行采样大大提高了性能。即使是非常相似的序列。如果在分析中使用超过100个分类群,则准确性和功效可以很高。我们提出以下建议:(1)正选择位点的预测对于少数密切相关的序列是不可行的;(2)使用大量的谱系是提高预测准确性和效力的最佳途径;(3)在真实的数据分析中应采用多个位点间异质选择压力模型。
Bayes prediction quantifies uncertainty by assigning posterior probabilities. It Was used to identify amino acids in a protein under recurrent diversifying selection indicated by higher nonsynonymous, (d(N)) than synonymous (d(S)) substitution rates or by omega = d(N)/d(S) > 1. Parameters were estimated by maximum likelihood under a codon substitution model that assumed several classes of sites with different w ratios. The Bayes theorem was used to calculate the posterior probabilities of each site falling into these site classes. Here. we evaluate the performance of Bayes prediction of amino acids under positive selection by computer simulation. We measured the accuracy by the proportion of predicted sites that were truly under selection and the power by the proportion of true positively selected sites that were predicted by the method. The accuracy was slightly better for longer sequences, whereas the power was largely unaffected by the increase in sequence length. Both accuracy and power were higher for medium or highly diverged sequences than for similar sequences. We found that accuracy and power were unacceptably low when data contained only a few highly similar sequences. However, sampling a large number of lineage improved the performance substantially. Even for very similar sequences. accuracy and Power can he high if over 100 taxa are used in the analysis. We make the following recommendations: (1) prediction of positive selection sites is not feasible for a few closely related sequences: (2) using it large number of lineages is the best way to improve the accuracy and power of the prediction: and (3) multiple models of heterogeneous selective pressures among sites should he applied in real data analysis.