Automatic assessment of alignment quality

Automatic assessment of alignment quality
复制标题

DOI:
10.1093/nar/gki1020
复制
发表时间:
2005-01-01
影响因子:
14.9
通讯作者:
Sonnhammer, ELL
Sonnhammer, ELL
中科院分区:
生物学2区
文献类型:
--
作者:
Lassmann, T;Sonnhammer, ELL

文献摘要

被引文献

相似文献

多序列比对在新基因组的注释中起着核心作用。鉴于这项任务的生物学和计算复杂性,自动生成高质量的对齐仍然具有挑战性。由于通常在数据分析管道的一开始就采用多个对齐,因此确保高对齐质量至关重要。我们描述了一个简单的,但优雅的,解决方案来评估生物校准的准确性自动。我们的方法是基于同一序列的几个比对的比较。我们引入了两个函数来比较比对:平均重叠分数和多重重叠分数。前者通过表达几个比对之间的相似性来识别困难的比对情况,而后者估计单个比对的生物学正确性。我们在MUMSA程序中实现了这两个函数,并在三个大型基准集上展示了这两个函数的总体鲁棒性和准确性。
Multiple sequence alignments play a central role in the annotation of novel genomes. Given the biological and computational complexity of this task, the automatic generation of high-quality alignments remains challenging. Since multiple alignments are usually employed at the very start of data analysis pipelines, it is crucial to ensure high alignment quality. We describe a simple, yet elegant, solution to assess the biological accuracy of alignments automatically. Our approach is based on the comparison of several alignments of the same sequences. We introduce two functions to compare alignments: the average overlap score and the multiple overlap score. The former identifies difficult alignment cases by expressing the similarity among several alignments, while the latter estimates the biological correctness of individual alignments. We implemented both functions in the MUMSA program and demonstrate the overall robustness and accuracy of both functions on three large benchmark sets.