Benchmarking spliced alignment programs including Spaln2, an extended version of Spaln that incorporates additional species-specific features.

Benchmarking spliced alignment programs including Spaln2, an extended version of Spaln that incorporates additional species-specific features.
复制标题

DOI:
10.1093/nar/gks708
复制
发表时间:
2012-11-01
影响因子:
14.9
通讯作者:
Gotoh O
Gotoh O
中科院分区:
生物学2区
文献类型:
--
作者:
Iwata H;Gotoh O

文献摘要

参考文献

被引文献

相似文献

剪接比对在真核基因结构的精确鉴定中起着核心作用。尽管已经开发了许多剪接比对程序,但最近 DNA 测序技术的快速进步需要软件工具的进一步改进。在各种条件下对算法进行基准测试是开发更好的软件必不可少的任务;然而,严重缺乏可用于对拼接对齐程序进行基准测试的适当数据集。在本研究中,我们构建了两种类型的数据集:模拟序列数据集和实际的跨物种数据集。数据集旨在对应各种实际情况,即不同的真核物种、不同类型的参考序列以及查询序列和目标序列之间的巨大差异。此外,我们还开发了程序 Spaln 的扩展版本,它在原始版本的评分方案中加入了两个附加功能,并根据我们的基准数据集检查了这个扩展版本 Spaln2 以及原始 Spaln 和其他代表性对齐器。尽管修改的效果并不单独显着,但 Spaln2 在大多数实际情况下始终是最准确且相当快的,特别是对于植物和真菌以及日益分歧的目标和查询序列对。
Spliced alignment plays a central role in the precise identification of eukaryotic gene structures. Even though many spliced alignment programs have been developed, recent rapid progress in DNA sequencing technologies demands further improvements in software tools. Benchmarking algorithms under various conditions is an indispensable task for the development of better software; however, there is a dire lack of appropriate datasets usable for benchmarking spliced alignment programs. In this study, we have constructed two types of datasets: simulated sequence datasets and actual cross-species datasets. The datasets are designed to correspond to various real situations, i.e. divergent eukaryotic species, different types of reference sequences, and the wide divergence between query and target sequences. In addition, we have developed an extended version of our program Spaln, which incorporates two additional features to the scoring scheme of the original version, and examined this extended version, Spaln2, together with the original Spaln and other representative aligners based on our benchmark datasets. Although the effects of the modifications are not individually striking, Spaln2 is consistently most accurate and reasonably fast in most practical cases, especially for plants and fungi and for increasingly divergent pairs of target and query sequences.
DOI: 10.1093/molbev/msp174
发表时间: 2009-11
影响因子: 10.7
作者:
Strope CL;Abel K;Scott SD;Moriyama EN
通讯作者: Moriyama EN
DOI: 10.1101/gr.6818908
发表时间: 2008-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Schwartz, Schraga H.;Silva, Joao;Ast, Gil
通讯作者: Ast, Gil
DOI: 10.1186/1471-2105-6-31
发表时间: 2005-02-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Slater GS;Birney E
通讯作者: Birney E
DOI: 10.1093/bioinformatics/16.3.190
发表时间: 2000-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Gotoh, O
通讯作者: Gotoh, O
DOI: 10.1093/bioinformatics/btp273
发表时间: 2009-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lu, David V.;Brown, Randall H.;Brent, Michael R.
通讯作者: Brent, Michael R.