Uncertainty in homology inferences: Assessing and improving genomic sequence alignment

Uncertainty in homology inferences: Assessing and improving genomic sequence alignment
复制标题

DOI:
10.1101/gr.6725608
复制
发表时间:
2008-02-01
期刊:
影响因子:
7
通讯作者:
Hein, Jotun
Hein, Jotun
中科院分区:
生物学1区
文献类型:
--
作者:
Lunter, Gerton;Rocco, Andrea;Hein, Jotun

文献摘要

被引文献

相似文献

序列比对是所有比较基因组学的基础,但它仍然是一个未完全解决的问题。特别是,推断的比对中的统计不确定性往往被忽视,而参数或系统发育推断被认为是毫无意义的没有信心的估计。在这里,我们报告的理论和模拟研究的基因组DNA成对比对在人类-小鼠的分歧。我们发现,在现有的全基因组比对中,> 15%的比对碱基是不正确的,并且我们确定了三种类型的比对错误,每种错误都导致所考虑的所有算法中的系统偏差。仔细建模的进化过程中提高对齐质量,然而,这些改进是适度的剩余的对齐错误相比,即使有确切的知识的进化模型,强调需要统计方法来考虑不确定性。我们开发了一种新的算法,边缘化后验解码(MPD),它明确占不确定性,是更少的偏见和更准确的比我们考虑的其他算法,并减少了三分之一的比例失调基地相比,现有的最好的算法。据我们所知,这是第一个非启发式算法的DNA序列比对,以显示强大的改进,在经典的Needleman-Wunsch算法。尽管如此,即使在改进的路线中也存在相当大的不确定性。我们的结论是,概率处理是必不可少的,以提高对齐质量和量化剩余的不确定性。随着人们越来越认识到非编码DNA的重要性,这一点变得越来越重要,非编码DNA的研究在很大程度上依赖于比对。校准误差是不可避免的,在从校准中得出结论时应予以考虑。在http://genserv.anat.ox.ac.uk/grape/上提供了帮助研究人员这样做的软件和比对。
Sequence alignment underpins all of comparative genomics, yet it remains an incompletely solved problem. In particular, the statistical uncertainty within inferred alignments is often disregarded, while parametric or phylogenetic inferences are considered meaningless without confidence estimates. Here, we report on a theoretical and simulation study of pairwise alignments of genomic DNA at human - mouse divergence. We find that > 15% of aligned bases are incorrect in existing whole- genome alignments, and we identify three types of alignment error, each leading to systematic biases in all algorithms considered. Careful modeling of the evolutionary process improves alignment quality; however, these improvements are modest compared with the remaining alignment errors, even with exact knowledge of the evolutionary model, emphasizing the need for statistical approaches to account for uncertainty. We develop a new algorithm, Marginalized Posterior Decoding ( MPD), which explicitly accounts for uncertainties, is less biased and more accurate than other algorithms we consider, and reduces the proportion of misaligned bases by a third compared with the best existing algorithm. To our knowledge, this is the first nonheuristic algorithm for DNA sequence alignment to show robust improvements over the classic Needleman - Wunsch algorithm. Despite this, considerable uncertainty remains even in the improved alignments. We conclude that a probabilistic treatment is essential, both to improve alignment quality and to quantify the remaining uncertainty. This is becoming increasingly relevant with the growing appreciation of the importance of noncoding DNA, whose study relies heavily on alignments. Alignment errors are inevitable, and should be considered when drawing conclusions from alignments. Software and alignments to assist researchers in doing this are provided at http://genserv.anat.ox.ac.uk/grape/.