Mind the gaps: Evidence of bias in estimates of multiple sequence alignments

Mind the gaps: Evidence of bias in estimates of multiple sequence alignments
复制标题

DOI:
10.1093/molbev/msm176
复制
发表时间:
2007-11-01
影响因子:
10.7
通讯作者:
Jermiin, Lars S.
Jermiin, Lars S.
中科院分区:
生物学1区
文献类型:
--
作者:
Golubchik, Tanya;Wise, Michael J.;Jermiin, Lars S.

文献摘要

被引文献

相似文献

多序列比对(MSA)是基因组和蛋白质组数据分析中至关重要的第一步。众所周知,常见的序列特征,如缺失和插入,会影响MSA程序的准确性,但尚未独立于序列变异的其他来源检查插入和缺失的位置对比对准确性的影响程度。我们评估了6个流行的MSA程序(ClustaW、DIALIGN-T、MAFFT、Muscle、PROBCONS和T-CAFICE)和一个实验程序恶作剧的性能,这些程序在氨基酸序列上只有一小段缺失残基的差异。分析表明,没有残基经常导致比对中间隙的错误放置,即使序列在其他方面是相同的。在包含部分重叠缺失的序列的数据集中,大多数MSA程序优先将缺口垂直比对,代价是不正确地比对侧翼区域中的残基。在评估的程序中,只有DIALIGN-T能够正确地相对于彼此放置重叠的间隙,但这通常取决于上下文,并且只在一些数据集中观察到。在包含非重叠缺失的序列的数据集中,DIALIGN-T和MAFFT(G-INS-I)都能够以近乎完美的精度比对缺口,但只有MAFFT一致地产生正确的比对。对于包含选择性剪接基因产物的异构体的数据集也是如此:DIALIGN-T和MAFFT都产生了高度准确的比对,MAFFT是这两个程序中更一致的。其他程序,特别是T-Cafee和CluastW,就不那么准确了。对于所有的数据集,不同的MSA程序产生的比对明显不同,表明依赖单一的MSA程序可能会产生误导的结果。因此,在处理可能包含缺失或插入的序列时,建议使用一个以上的MSA程序,特别是对于高通量和流水线应用,其中人工精炼每个比对是不可行的。
Multiple sequence alignment (MSA) is a crucial first step in the analysis of genomic and proteomic data. Commonly occurring sequence features, such as deletions and insertions, are known to affect the accuracy of MSA programs, but the extent to which alignment accuracy is affected by the positions of insertions and deletions has not been examined independently of other sources of sequence variation. We assessed the performance of 6 popular MSA programs (ClustalW, DIALIGN-T, MAFFT, MUSCLE, PROBCONS, and T-COFFEE) and one experimental program, PRANK, on amino acid sequences that differed only by short regions of deleted residues. The analysis showed that the absence of residues often led to an incorrect placement of gaps in the alignments, even though the sequences were otherwise identical. In data sets containing sequences with partially overlapping deletions, most MSA programs preferentially aligned the gaps vertically at the expense of incorrectly aligning residues in the flanking regions. Of the programs assessed, only DIALIGN-T was able to place overlapping gaps correctly relative to one another, but this was usually context dependent and was observed only in some of the data sets. In data sets containing sequences with non-overlapping deletions, both DIALIGN-T and MAFFT (G-INS-I) were able to align gaps with near-perfect accuracy, but only MAFFT produced the correct alignment consistently. The same was true for data sets that comprised isoforms of alternatively spliced gene products: both DIALIGN-T and MAFFT produced highly accurate alignments, with MAFFT being the more consistent of the 2 programs. Other programs, notably T-COFFEE and ClustalW, were less accurate. For all data sets, alignments produced by different MSA programs differed markedly, indicating that reliance on a single MSA program may give misleading results. It is therefore advisable to use more than one MSA program when dealing with sequences that may contain deletions or insertions, particularly for high-throughput and pipeline applications where manual refinement of each alignment is not practicable.