An improved general amino acid replacement matrix

An improved general amino acid replacement matrix
复制标题

DOI:
10.1093/molbev/msn067
复制
发表时间:
2008-07-01
影响因子:
10.7
通讯作者:
Gascuel, Olivier
Gascuel, Olivier
中科院分区:
生物学1区
文献类型:
--
作者:
Le, Si Quang;Gascuel, Olivier

文献摘要

被引文献

相似文献

氨基酸替代矩阵是蛋白质系统发育学的重要基础。它们用于计算沿系统发育分支的替换概率,从而计算数据的可能性。它们对于蛋白质排列也至关重要。自 Dayhoff 等人的开创性工作以来,已经提出了许多替换矩阵和根据蛋白质比对估计这些矩阵的方法。 (1972)。 Whelan 和 Goldman (2001) 及其 WAG 矩阵取得了重要进展,这要归功于有效的最大似然估计方法,该方法考虑了每个训练比对中序列的系统发育。我们通过在矩阵估计中纳入跨位点进化速率的变异性并使用比用于估计 WAG 的 BRKALN 更大且多样化的数据库来进一步完善该方法。为了估计我们的新矩阵(以作者的名字命名为 LG),我们使用 XRATE 软件的改编版和 Pfam 的 3,912 个比对,其中包含大约 50,000 个序列和大约 650 万个残基。为了评估 LG 性能,我们使用由 TreeBase 中的 59 个比对组成的独立样本,并将 Pfam 比对随机分为 3,412 个训练比对和 500 个测试比对。与 WAG 和 JTT 的比较显示出明显的可能性改进。通过 TreeBase,我们发现 1)与 WAG 和 JTT 相比,每个站点的平均 Akaike 信息标准增益分别为 0.25 和 0.42; 2)LG 在 38 个比对(共 59 个)中显着优于 WAG,仅在 2 个比对中显着较差; 3)使用 LG、WAG 和 JTT 推断的树拓扑经常不同,这表明使用 LG 不仅影响似然值,而且影响输出树。 Pfam 的测试比对结果是类似的。 LG 和 PHYML 实现可以从 http://atgc.lirmm.fr/LG 下载。
Amino acid replacement matrices are an essential basis of protein phylogenetics. They are used to compute substitution probabilities along phylogeny branches and thus the likelihood of the data. They are also essential in protein alignment. A number of replacement matrices and methods to estimate these matrices from protein alignments have been proposed since the seminal work of Dayhoff et al. (1972). An important advance was achieved by Whelan and Goldman (2001) and their WAG matrix, thanks to an efficient maximum likelihood estimation approach that accounts for the phylogenies of sequences within each training alignment. We further refine this method by incorporating the variability of evolutionary rates across sites in the matrix estimation and using a much larger and diverse database than BRKALN, which was used to estimate WAG. To estimate our new matrix (called LG after the authors), we use an adaptation of the XRATE software and 3,912 alignments from Pfam, comprising similar to 50,000 sequences and similar to 6.5 million residues overall. To evaluate the LG performance, we use an independent sample consisting of 59 alignments from TreeBase and randomly divide Pfam alignments into 3,412 training and 500 test alignments. The comparison with WAG and JTT shows a clear likelihood improvement. With TreeBase, we find that 1) the average Akaike information criterion gain per site is 0.25 and 0.42, when compared with WAG and JTT, respectively; 2) LG is significantly better than WAG for 38 alignments (among 59), and significantly worse with 2 alignments only; and 3) tree topologies inferred with LG, WAG, and JTT frequently differ, indicating that using LG impacts not only the likelihood value but also the output tree. Results with the test alignments from Pfam are analogous. LG and a PHYML implementation can be downloaded from http://atgc.lirmm.fr/LG.