Towards a reliable objective function for multiple sequence alignments

Towards a reliable objective function for multiple sequence alignments
复制标题

DOI:
10.1006/jmbi.2001.5187
复制
发表时间:
2001-12-07
影响因子:
5.6
通讯作者:
Poch, O
Poch, O
中科院分区:
生物学2区
文献类型:
--
作者:
Thompson, JD;Plewniak, F;Poch, O

文献摘要

被引文献

相似文献

多序列比对是现代分子生物学中许多不同领域的基本工具,包括蛋白质家族的功能和进化研究,多序列比对在新的基因组注释和分析集成系统中也扮演着重要的角色。因此,本着为数据库搜索技术评估成对序列比对的工作精神,开发新的多重比对分数和统计数据是至关重要的。本文提出了一种新的多序列比对客观评分函数--NorMD。NorMD结合了柱评分技术的优点和结合残基相似性分数的方法的敏感性。此外,norMD包含从头算序列信息,例如要比对的序列的数量、长度和相似性。使用SCOP和BAliBASE数据库中的结构比对证明了norMD目标函数的敏感性和可靠性。NorMD分数随后被应用于BlastP检测到的完整序列(MAC)的多重比对,用于由霍乱弧菌基因组编码的一组734个假设蛋白质的E-Value<10。不相关或比对不良的序列被自动从MAC中移除,留下高质量的多重比对,可以在随后的功能和/或结构注释过程中可靠地利用这些比对。在去除不可靠序列后,176个(24%)的比对包含至少一个带有功能注释的序列。其中103个新的匹配得到了Interpro域和Motif数据库的重大匹配的支持。(C)2001年学术出版社。
Multiple sequence alignment is a fundamental tool in a number of different domains in modem molecular biology, including functional and evolutionary studies of a protein family, Multiple alignments also play an essential role in the new integrated systems for genome annotation and analysis. Thus, the development of new multiple alignment scores and statistics is essential, in the spirit of the work dedicated to the evaluation of pairwise sequence alignments for database searching techniques. We present here norMD, a new objective scoring function for multiple sequence alignments. NorMD combines the advantages of the column-scoring techniques with the sensitivity of methods incorporating residue similarity scores. In addition, norMD incorporates ab initio sequence information, such as the number, length and similarity of the sequences to be aligned. The sensitivity and reliability of the norMD objective function is demonstrated using structural alignments in the SCOP and BAliBASE databases. The norMD scores are then applied to the multiple alignments of the complete sequences (MACS) detected by BlastP with E-value < 10, for a set of 734 hypothetical proteins encoded by the Vibrio cholerae genome. Unrelated or badly aligned sequences were automatically removed from the MACS, leaving a high-quality multiple alignment which could be reliably exploited in a subsequent functional and/or structural annotation process. After removal of unreliable sequences, 176 (24%) of the alignments contained at least one sequence with a functional annotation. 103 of these new matches were supported by significant hits to the Interpro domain and motif database. (C) 2001 Academic Press.