SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures

SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures
复制标题

DOI:
10.1093/bioinformatics/bti582
复制
发表时间:
2005-09-15
期刊:
影响因子:
5.8
通讯作者:
Zhou, YQ
Zhou, YQ
中科院分区:
生物学3区
文献类型:
--
作者:
Zhou, HY;Zhou, YQ

文献摘要

被引文献

相似文献

动机:多序列比对是基因及其进化关系的基因组尺度研究的生物信息学工具的重要组成部分。然而,在远程同源物之间进行精确的比对是具有挑战性的。在这里,我们开发了一种称为SPEM的方法,该方法使用预处理的序列剖面和预测的二级结构进行两两比对,基于一致性的评分用于两两比对的细化,并使用渐进算法进行最终的多重比对。结果:spm的比对精度与现有的ClustalW、T-Coffee、MUSCLE、ProbCons和PRALINE(PSI)等方法在易(同源)和难(远程同源)基准上进行了比较。结果表明,SPEM方法的比对结果平均比对分数比其他方法高7 ~ 15%(序列一致性< 30%)。它对同源物比对的准确性(序列同一性bbb30 %)在统计上与最先进的技术(如ProbCons或MUSCLE 6.0)没有区别。
Motivation: Multiple sequence alignment is an essential part of bioinformatics tools for a genome-scale study of genes and their evolution relations. However, making an accurate alignment between remote homologs is challenging. Here, we develop a method, called SPEM, that aligns multiple sequences using pre-processed sequence profiles and predicted secondary structures for pairwise alignment, consistency-based scoring for refinement of the pairwise alignment and a progressive algorithm for final multiple alignment.Results: The alignment accuracy of SPEM is compared with those of established methods such as ClustalW, T-Coffee, MUSCLE, ProbCons and PRALINE(PSI) in easy (homologs) and hard (remote homologs) benchmarks. Results indicate that the average sum of pairwise alignment scores given by SPEM are 7-15% higher than those of the methods compared in aligning remote homologs (sequence identity < 30%). Its accuracy for aligning homologs (sequence identity > 30%) is statistically indistinguishable from those of the state-of-the-art techniques such as ProbCons or MUSCLE 6.0.