A collection of amino acid replacement matrices derived from clusters of orthologs.

A collection of amino acid replacement matrices derived from clusters of orthologs.
复制标题

源自直系同源簇的氨基酸置换矩阵的集合。

DOI:
10.1007/s00239-005-0060-0
复制
发表时间:
2005
影响因子:
3.9
通讯作者:
Loomis,WilliamF
Loomis,WilliamF
中科院分区:
生物学3区
文献类型:
--
作者:
Olsen,Rolf;Loomis,WilliamF

文献摘要

相似文献

利用34个氨基酸替代矩阵、序列上下文分析和系统发育树对同源蛋白之间的序列差异进行了表征。该模型是在15种生物(包括原生生物、植物、盘基ostelium、真菌和动物)中排列的蛋白质序列的大型数据集上进行训练的。与目前用于系统发育的模型(即JTT+Γ±F和WAG+Γ±F)进行的比较测试表明,我们的模型在与测试数据集相似的数据集上应该优于JTT+Γ±F和WAG+Γ±F模型。测试数据集包含来自上述所有五个主要分类类群的蛋白质序列的380个多个比对。我们的同源蛋白序列发散模型的强大性能可归因于它能够更好地近似氨基酸平衡频率到比对柱中发现的成分。
Sequence divergence among orthologous proteins was characterized with 34 amino acid replacement matrices, sequence context analysis, and a phylogenetic tree. The model was trained on very large datasets of aligned protein sequences drawn from 15 organisms including protists, plants,Dictyostelium, fungi, and animals. Comparative tests with models currently used in phylogeny, i.e., with JTT+Γ±F and WAG+Γ±F, made on a test dataset of 380 multiple alignments containing protein sequences from all five of the major taxonomic groups mentioned, indicate that our model should be preferred over the JTT+Γ±F and WAG+Γ±F models on datasets similar to the test dataset. The strong performance of our model of orthologous protein sequence divergence can be attributed to its ability to better approximate amino acid equilibrium frequencies to compositions found in alignment columns.