Accurate statistical model of comparison between multiple sequence alignments.

Accurate statistical model of comparison between multiple sequence alignments.
复制标题

DOI:
10.1093/nar/gkn065
复制
发表时间:
2008-04
影响因子:
14.9
通讯作者:
Grishin NV
Grishin NV
中科院分区:
生物学2区
文献类型:
--
作者:
Sadreyev RI;Grishin NV

文献摘要

参考文献

被引文献

相似文献

多种蛋白质序列比对(MSA)的比较揭示了蛋白质家族之间意外的进化关系,并导致对空间结构和功能的令人兴奋的预测。 MSA比较的能力严重取决于用于对数据库搜索中发现的相似性进行排列的统计模型的质量,以便将生物学相关的关系与虚假连接区分开。在这里,我们开发了MSA比较的准确统计描述,该描述不是源于单个序列比较的常规模型,并捕获了蛋白质家族的基本特征。作为最终结果,我们使用取决于MSA长度和序列多样性的数学函数来计算任何两个MSA之间的相似性。为了开发这些统计显着性的估计值,我们首先建立了一种生成逼真的诱饵的程序,该诱饵重现了由蛋白质二级结构决定的序列保护的自然模式。其次,由于这些比对之间的相似性得分不遵循经典的牙龈极端价值分布,因此我们提出了一个新颖的分布,该分布与数据具有统计上完美的一致性。第三,我们将此随机模型应用于数据库搜索,并表明它超过了传统模型,以检测远程蛋白质相似性的准确性。
Comparison of multiple protein sequence alignments (MSA) reveals unexpected evolutionary relations between protein families and leads to exciting predictions of spatial structure and function. The power of MSA comparison critically depends on the quality of statistical model used to rank the similarities found in a database search, so that biologically relevant relationships are discriminated from spurious connections. Here, we develop an accurate statistical description of MSA comparison that does not originate from conventional models of single sequence comparison and captures essential features of protein families. As a final result, we compute E-values for the similarity between any two MSA using a mathematical function that depends on MSA lengths and sequence diversity. To develop these estimates of statistical significance, we first establish a procedure for generating realistic alignment decoys that reproduce natural patterns of sequence conservation dictated by protein secondary structure. Second, since similarity scores between these alignments do not follow the classic Gumbel extreme value distribution, we propose a novel distribution that yields statistically perfect agreement with the data. Third, we apply this random model to database searches and show that it surpasses conventional models in the accuracy of detecting remote protein similarities.
DOI: 10.1016/s0097-8485(96)80004-0
发表时间: 1996-03-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gribskov, M;Robinson, NL
通讯作者: Robinson, NL
Pfam:氏族、网络工具和服务。
DOI: 10.1093/nar/gkj149
发表时间: 2006-01-01
影响因子: 14.9
作者:
Finn, Robert D.;Mistry, Jaina;Schuster-Bockler, Benjamin;Griffiths-Jones, Sam;Hollich, Volker;Lassmann, Timo;Moxon, Simon;Marshall, Mhairi;Khanna, Ajay;Durbin, Richard;Eddy, Sean R.;Sonnhammer, Erik L. L.;Bateman, Alex
通讯作者: Bateman, Alex
DOI: 10.1002/prot.20184
发表时间: 2004-10-01
影响因子: 2.9
作者:
Ohlson, T;Wallner, B;Elofsson, A
通讯作者: Elofsson, A
DOI: 10.1093/nar/gkh370
发表时间: 2004-07-01
影响因子: 14.9
作者:
Ginalski, K;von Grotthuss, M;Rychlewski, L
通讯作者: Rychlewski, L
DOI: 10.1093/nar/gkg504
发表时间: 2003-07-01
影响因子: 14.9
作者:
Ginalski, K;Pas, J;Rychlewski, L
通讯作者: Rychlewski, L