A Bayesian mixture model for across-site heterogeneities in the amino-acid replacement process

A Bayesian mixture model for across-site heterogeneities in the amino-acid replacement process
复制标题

DOI:
10.1093/molbev/msh112
复制
发表时间:
2004-06-01
影响因子:
10.7
通讯作者:
Philippe, H
Philippe, H
中科院分区:
生物学1区
文献类型:
--
作者:
Lartillot, N;Philippe, H

文献摘要

被引文献

相似文献

大多数目前的序列进化模型假设蛋白质的所有位点在相同的替换过程下进化,其特征在于20 X 20替换矩阵。在这里,我们建议放松这一假设,通过开发贝叶斯混合模型,允许在不同位点的蛋白质比对的氨基酸替换模式被描述由不同的取代过程。我们的模型,名为CAT,假设存在不同的过程(或类)不同的平衡频率超过20个残基。通过使用狄利克雷过程先验,类的总数和它们各自的氨基酸谱,以及每个位点与给定类的隶属关系,都是模型的自由变量。通过这种方式,CAT模型能够适应数据中实际存在的复杂性,并通过类的后验平均数来估计替代异质性。我们发现,一个显着的异质性水平是存在于蛋白质的取代模式,而标准的一个矩阵模型无法解释这种异质性。通过评估贝叶斯因子,我们证明了CAT在我们分析的所有数据集上都优于标准模型。总之,这些结果表明,取代的真实的序列的模式的复杂性,更好地捕捉CAT模型,提供了可能性,研究其影响系统发育重建及其与结构-功能决定因素的连接。
Most current models of sequence evolution assume that all sites of a protein evolve under the same substitution process, characterized by a 20 X 20 substitution matrix. Here, we propose to relax this assumption by developing a Bayesian mixture model that allows the amino-acid replacement pattern at different sites of a protein alignment to be described by distinct substitution processes. Our model, named CAT, assumes the existence of distinct processes (or classes) differing by their equilibrium frequencies over the 20 residues. Through the use of a Dirichlet process prior, the total number of classes and their respective amino-acid profiles, as well as the affiliations of each site to a given class, are all free variables of the model. In this way, the CAT model is able to adapt to the complexity actually present in the data, and it yields an estimate of the substitutional heterogeneity through the posterior mean number of classes. We show that a significant level of heterogeneity is present in the substitution patterns of proteins, and that the standard one-matrix model fails to account for this heterogeneity. By evaluating the Bayes factor, we demonstrate that the standard model is outperformed by CAT on all of the data sets which we analyzed. Altogether, these results suggest that the complexity of the pattern of substitution of real sequences is better captured by the CAT model, offering the possibility of studying its impact on phylogenetic reconstruction and its connections with structure-function determinants.