Direct mapping and alignment of protein sequences onto genomic sequence

Direct mapping and alignment of protein sequences onto genomic sequence
复制标题

DOI:
10.1093/bioinformatics/btn460
复制
发表时间:
2008-11-01
期刊:
影响因子:
5.8
通讯作者:
Gotoh, Osamu
Gotoh, Osamu
中科院分区:
生物学3区
文献类型:
--
作者:
Gotoh, Osamu

文献摘要

被引文献

相似文献

动机:在新确定的基因组序列中寻找蛋白质编码基因是了解基因组中写入内容的第一步。与没有此类知识的情况相比,同源基因的转录本序列(如果可用)可以显着提高基因及其结构预测的准确性。由于蛋白质序列通常比核苷酸序列更保守,因此可以使用远程同源物作为模板,扩展了基于证据的基因识别方法的适用性。然而,到目前为止,似乎还没有开发出能够在哺乳动物大小的基因组序列上同时绘制和比对多个蛋白质序列的工具。结果:我们已经扩展了我们的计算机程序 Spaln 以接受蛋白质序列以及 cDNA 序列作为查询。当查询和目标序列相当相似时,例如在哺乳动物直向同源物之间,Spaln 的运行速度比传统方法快一到两个数量级,传统方法依赖于 Blast 搜索,然后进行基于动态编程的剪接比对。 Spaln 的外显子级和基因级准确率明显高于同类最佳可用方法所获得的准确率,特别是当查询和目标关系较远时。
Motivation: Finding protein-coding genes in a newly determined genomic sequence is the first step toward understanding the content written in the genome. Sequences of transcripts of homologous genes, if available, can considerably improve accuracy of prediction of genes and their structures, compared with that without such knowledge. As protein sequences are generally better conserved than nucleotide sequences, remote homologs can be used as templates, extending the applicability of evidence-based gene recognition methods. However, no tool seems to have been developed so far to simultaneously map and align a number of protein sequences on mammalian-sized genomic sequence.Results: We have extended our computer program Spaln to accept protein sequences, as well as cDNA sequences, as queries. When the query and the target sequences are reasonably similar, e.g. between mammalian orthologs, Spaln runs one to two orders of magnitude faster than conventional approaches that rely on Blast search followed by dynamic-programming-based spliced alignment. Exon-level and gene-level accuracies of Spaln are significantly higher than those obtained by the best available methods of the same type, particularly when the query and the target are distantly related.