An enhanced RNA alignment benchmark for sequence alignment programs.

An enhanced RNA alignment benchmark for sequence alignment programs.
复制标题

DOI:
10.1186/1748-7188-1-19
复制
发表时间:
2006-10-24
期刊:
Algorithms for molecular biology : AMB
影响因子:
--
通讯作者:
Steger G
Steger G
中科院分区:
其他
文献类型:
--
作者:
Wilm A;Mainz I;Steger G

文献摘要

参考文献

被引文献

相似文献

比对程序的性能传统上是在蛋白质序列集上进行测试的,其中参考比对是已知的。从这样的蛋白质基准得出的结论不一定适用于RNA比对问题,正如迄今为止发表的第一个RNA比对基准所证明的那样。例如,黄光区--比对质量急剧下降的相似性范围--RNA的起始值为60%,而蛋白质的起始值为20%。在这项研究中,我们对以前的基准进行了改进。基准数据库中的RNA序列集取自数量增加的RNA家族,以避免仅使用几个家族造成的意外影响。集合的大小从2到15个序列不等,以评估序列数量对程序性能的影响。比对质量由两个衡量标准来评分:一个只考虑核苷酸匹配,另一个衡量结构保守。参数的性能顺序--如核苷酸替代矩阵和缺口成本--以及程序的性能顺序通过等级检验进行评级。大多数序列比对程序在具有高序列同一性的RNA序列集上执行得同样好,即具有75%以上的平均成对序列同一性(APSI)。间隙张开和间隙延伸的参数对配准质量的影响较大,低于APSI≤的75%;给出了几种方案的最优参数组合。使用不同的4×4替换矩阵仅在某些情况下提高了程序性能。随着序列号的增加和/或序列一致性的减少,迭代程序的性能显著提高,这使得它们明显优于使用纯非迭代、渐进方法的程序。最好的序列比对程序产生高质量的比对,低到APSI;55%;在APSI较低的情况下,建议使用序列+结构比对程序。
The performance of alignment programs is traditionally tested on sets of protein sequences, of which a reference alignment is known. Conclusions drawn from such protein benchmarks do not necessarily hold for the RNA alignment problem, as was demonstrated in the first RNA alignment benchmark published so far. For example, the twilight zone – the similarity range where alignment quality drops drastically – starts at 60 % for RNAs in comparison to 20 % for proteins. In this study we enhance the previous benchmark. The RNA sequence sets in the benchmark database are taken from an increased number of RNA families to avoid unintended impact by using only a few families. The size of sets varies from 2 to 15 sequences to assess the influence of the number of sequences on program performance. Alignment quality is scored by two measures: one takes into account only nucleotide matches, the other measures structural conservation. The performance order of parameters – like nucleotide substitution matrices and gap-costs – as well as of programs is rated by rank tests. Most sequence alignment programs perform equally well on RNA sequence sets with high sequence identity, that is with an average pairwise sequence identity (APSI) above 75 %. Parameters for gap-open and gap-extension have a large influence on alignment quality lower than APSI ≤ 75 %; optimal parameter combinations are shown for several programs. The use of different 4 × 4 substitution matrices improved program performance only in some cases. The performance of iterative programs drastically increases with increasing sequence numbers and/or decreasing sequence identity, which makes them clearly superior to programs using a purely non-iterative, progressive approach. The best sequence alignment programs produce alignments of high quality down to APSI > 55 %; at lower APSI the use of sequence+structure alignment programs is recommended.
DOI: 10.1007/bf00818163
发表时间: 1994-02-01
影响因子: 1.8
作者:
HOFACKER, IL;FONTANA, W;SCHUSTER, P
通讯作者: SCHUSTER, P
DOI: 10.1093/nar/gkg006
发表时间: 2003-01-01
影响因子: 14.9
作者:
Griffiths-Jones, S;Bateman, A;Eddy, SR
通讯作者: Eddy, SR
DOI: 10.1093/nar/gkg500
发表时间: 2003-07-01
影响因子: 14.9
作者:
Chenna, R;Sugawara, H;Thompson, JD
通讯作者: Thompson, JD
DOI: 10.1006/jmbi.1996.0679
发表时间: 1996-12-13
影响因子: 5.6
作者:
Gotoh, O
通讯作者: Gotoh, O
DOI: 10.1093/bioinformatics/bti279
发表时间: 2005-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Havgaard, JH;Lyngso, RB;Gorodkin, J
通讯作者: Gorodkin, J