Pairwise sequence alignment below the twilight zone

Pairwise sequence alignment below the twilight zone
复制标题

DOI:
10.1006/jmbi.2001.4495
复制
发表时间:
2001-03-23
影响因子:
5.6
通讯作者:
Cohen, FE
Cohen, FE
中科院分区:
生物学2区
文献类型:
--
作者:
Blake, JD;Cohen, FE

文献摘要

被引文献

相似文献

在低对同源性下改进序列比对对于在数据库搜索中识别潜在的远程同源物以及获得准确的序列比对作为同源性建模的前奏是非常重要的。我们的工作有两个动机:结构数据为开发技术提供了良好的训练示例,以改善远程同源物的比对;远端同源物的一般替代模式与密切相关的蛋白质不同。本文介绍了一套基于结构叠加数据的氨基酸残基交换矩阵。这些矩阵利用已知的结构同源性作为表征进化对残基取代谱的影响的手段。考虑到它们的来源,单个残基-残基交换频率在化学上是敏感的就不足为奇了。结果表明,结构交换矩阵在两两比对精度和跨远缘序列的功能注释/折叠识别精度上均有显著提高。我们通过使用从结构数据库中提取的同源结构域的叠加作为金标准,证明了改进的成对比对,并继续显示使用同源折叠家族数据库提高了折叠识别的准确性。将该方法应用于幽门螺杆菌基因组中未分配的开放阅读框,以确定5个匹配,其中2个在序列数据库中没有新的注释。此外,我们描述了一个新的循环排列策略,以确定远同源经历基因复制和随后的缺失。使用这种方法,我们已经从幽门螺杆菌基因组中确定了一个潜在的同源物,该同源物与先前未分配的开放阅读框相同。(C) 2001学术出版社。
Improved sequence alignment at low pairwise identity is important for identifying potential remote homologues in database searches and for obtaining accurate alignments as a prelude to modeling structures by homology. Our work is motivated by two observations: structural data provide superior training examples for developing techniques to improve the alignment of remote homologues; and general substitution patterns for remote homologues differ from those of closely related proteins. We introduce a new set of amino acid residue interchange matrices built from structural superposition data. These matrices exploit known structural homology as a means of characterizing the effect evolution has on residue-substitution profiles. Given their origin, it is not surprising that the individual residue-residue interchange frequencies are chemically sensible.The structural interchange matrices show a significant increase both in pairwise alignment accuracy and in functional annotation/fold recognition accuracy across distantly related sequences. We demonstrate improved pairwise alignment by using superpositions of homologous domains extracted from a structural database as a gold standard and go on to show an increase in fold recognition accuracy using a database of homologous fold families. This was applied to the unassigned open reading frames from the genome of Helicobacter pylori to identify five matches, two of which are not represented by new annotations in the sequence databases. In addition, we describe a new cyclic permutation strategy to identify distant homologues that experienced gene duplication and subsequent deletions. Using this method, we have identified a potential homologue to one additional previously unassigned open reading frame from the H. pylori genome. (C) 2001 Academic Press.