Splign: algorithms for computing spliced alignments with identification of paralogs.

Splign: algorithms for computing spliced alignments with identification of paralogs.
复制标题

DOI:
10.1186/1745-6150-3-20
复制
发表时间:
2008-05-21
期刊:
影响因子:
5.5
通讯作者:
Lipman D
Lipman D
中科院分区:
生物学2区
文献类型:
--
作者:
Kapustin Y;Souvorov A;Tatusova T;Lipman D

文献摘要

参考文献

被引文献

相似文献

计算 cDNA 序列与基因组的精确比对是现代基因组注释流程的基础。旁系同源物、小外显子、非一致性剪接信号、测序错误和多态性位点的存在等几个因素给现有剪接比对算法带来了公认的困难。我们描述了名为 Splign 的工具背后的一组算法,用于计算 cDNA 到基因组的比对。该算法包括高性能初步比对、基于相邻重复区域的正式定义模型的区室识别以及精细的序列比对。在一系列测试中,Splign 在合理的时间内产生了比其他常用于计算拼接比对的工具更准确的结果。 Splign 能够处理使剪接比对问题复杂化的各种问题,这使其成为真核基因组注释过程和选择性剪接研究中的有用工具。其性能足以在几个小时内对齐当前可用的最大 cDNA 数据池,例如在中等规模的计算集群上设置的人类 EST。重复识别(区室化)算法可以独立用于其他领域,例如假基因的研究。本文审阅者:Steven Salzberg、Arcady Mushegian 和 Andrey Mironov(由 Mikhail Gelfand 提名)。
The computation of accurate alignments of cDNA sequences against a genome is at the foundation of modern genome annotation pipelines. Several factors such as presence of paralogs, small exons, non-consensus splice signals, sequencing errors and polymorphic sites pose recognized difficulties to existing spliced alignment algorithms. We describe a set of algorithms behind a tool called Splign for computing cDNA-to-Genome alignments. The algorithms include a high-performance preliminary alignment, a compartment identification based on a formally defined model of adjacent duplicated regions, and a refined sequence alignment. In a series of tests, Splign has produced more accurate results than other tools commonly used to compute spliced alignments, in a reasonable amount of time. Splign's ability to deal with various issues complicating the spliced alignment problem makes it a helpful tool in eukaryotic genome annotation processes and alternative splicing studies. Its performance is enough to align the largest currently available pools of cDNA data such as the human EST set on a moderate-sized computing cluster in a matter of hours. The duplications identification (compartmentization) algorithm can be used independently in other areas such as the study of pseudogenes. This article was reviewed by: Steven Salzberg, Arcady Mushegian and Andrey Mironov (nominated by Mikhail Gelfand).
DOI: 10.1089/10665270050081478
发表时间: 2000-02-01
影响因子: 1.7
作者:
Zhang, Z;Schwartz, S;Miller, W
通讯作者: Miller, W
DOI: 10.1101/gr.194201
发表时间: 2001-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Ning, ZM;Cox, AJ;Mullikin, JC
通讯作者: Mullikin, JC
DOI: 10.1101/gr.195301
发表时间: 2001-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Wheelan, SJ;Church, DM;Ostell, JM
通讯作者: Ostell, JM
DOI: 10.1093/bioinformatics/bti310
发表时间: 2005-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Wu, TD;Watanabe, CK
通讯作者: Watanabe, CK
DOI: 10.1093/bioinformatics/bti774
发表时间: 2006-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Morgulis, A;Gertz, EM;Agarwala, R
通讯作者: Agarwala, R