Exon discovery by genomic sequence alignment

Exon discovery by genomic sequence alignment
复制标题

DOI:
10.1093/bioinformatics/18.6.777
复制
发表时间:
2002-06-01
期刊:
影响因子:
5.8
通讯作者:
Mewes, HW
Mewes, HW
中科院分区:
生物学3区
文献类型:
--
作者:
Morgenstern, B;Rinner, O;Mewes, HW

文献摘要

被引文献

相似文献

动机:在进化过程中,基因组序列中的功能区域往往比随机突变的“垃圾DNA”更保守,因此局部序列相似性通常表明生物功能。这一事实可用于通过跨物种序列比较来鉴定大型真核生物DNA序列中的功能元件。近年来,已经提出了几种基因预测方法,通过比较匿名的基因组序列,例如人类和小鼠的基因组序列。这些方法的主要优点是它们基于简单且普遍适用的(局部)序列相似性度量;与标准的基因发现方法不同,它们不依赖于物种特异性训练数据或数据库中同源基因的存在。与所有比较序列分析方法一样,新的比较基因发现方法严重依赖于潜在序列比对的质量。结果:在这里,我们描述了一种新的序列比对程序DIALIGN的实现,该程序已开发用于大型基因组序列的比对。我们将我们的方法与比对程序PipMaker、WABA和BLAST进行了比较,我们发现这些程序识别的局部相似性与蛋白质编码区域高度相关。在我们的测试中,PipMaker是最敏感的方法,而DIALIGN是最具体的方法。
Motivation: During evolution, functional regions in genomic sequences tend to be more highly conserved than randomly mutating 'junk DNA' so local sequence similarity often indicates biological functionality. This fact can be used to identify functional elements in large eukaryotic DNA sequences by cross-species sequence comparison. In recent years, several gene-prediction methods have been proposed that work by comparing anonymous genomic sequences, for example from human and mouse. The main advantage of these methods is that they are based on simple and generally applicable measures of (local) sequence similarity; unlike standard gene-finding approaches they do not depend on species-specific training data or on the presence of cognate genes in data bases. As all comparative sequence-analysis methods, the new comparative gene-finding approaches critically rely on the quality of the underlying sequence alignments.Results: Herein, we describe a new implementation of the sequence-alignment program DIALIGN that has been developed for alignment of large genomic sequences. We compare our method to the alignment programs PipMaker, WABA and BLAST and we show that local similarities identified by these programs are highly correlated to protein-coding regions. In our test runs, PipMaker was the most sensitive method while DIALIGN was most specific.