SPARSE: quadratic time simultaneous alignment and folding of RNAs without sequence-based heuristics.

SPARSE: quadratic time simultaneous alignment and folding of RNAs without sequence-based heuristics.
复制标题

DOI:
10.1093/bioinformatics/btv185
复制
发表时间:
2015-08-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Backofen R
Backofen R
中科院分区:
其他
文献类型:
--
作者:
Will S;Otto C;Miladi M;Möhl M;Backofen R

文献摘要

被引文献

相似文献

动机:RNA-Seq实验揭示了大量新的ncrna。基于同时对齐和折叠的分析的黄金标准受到极端时间复杂性的影响。随后,许多更快的“sankoff风格”方法被提出。通常,这些方法的性能依赖于基于序列的启发式算法,该算法将搜索空间限制在最优或接近最优的序列比对上;然而,基于序列的方法的准确性在序列身份低于60%的rna中崩溃。像LocARNA这样不需要基于序列的启发式方法的对齐方法,一直被限制在高复杂性(四次时间)。结果:我们打破了这一障碍,引入了新的sankoff式算法“基于rna结构集成(SPARSE)的稀疏预测和对齐”,该算法在二次时间内运行,不需要基于序列的启发式算法。为了实现这种低复杂度,与序列比对算法相当,SPARSE基于RNA集成的结构特性进行了强稀疏化。继PMcomp之后,稀疏从轻量级能量计算中获得了进一步的加速。尽管所有现有的轻量级Sankoff风格方法都通过禁止循环删除和插入来限制Sankoff的原始模型,但SPARSE首次将Sankoff算法完全转移到轻量级能量模型中。与LocARNA相比,SPARSE在更短的时间内实现了相似的对齐和更好的折叠质量(加速:3.7)。在类似的运行时,它比使用基于序列的启发式方法的RAF更准确地对齐低序列标识实例。可用性和实现:在http://www.bioinf.uni-freiburg.de/Software/SPARSE上可以免费获得SPARSE。补充信息:补充数据可在Bioinformatics在线获取。
Motivation: RNA-Seq experiments have revealed a multitude of novel ncRNAs. The gold standard for their analysis based on simultaneous alignment and folding suffers from extreme time complexity of . Subsequently, numerous faster ‘Sankoff-style’ approaches have been suggested. Commonly, the performance of such methods relies on sequence-based heuristics that restrict the search space to optimal or near-optimal sequence alignments; however, the accuracy of sequence-based methods breaks down for RNAs with sequence identities below 60%. Alignment approaches like LocARNA that do not require sequence-based heuristics, have been limited to high complexity ( quartic time). Results: Breaking this barrier, we introduce the novel Sankoff-style algorithm ‘sparsified prediction and alignment of RNAs based on their structure ensembles (SPARSE)’, which runs in quadratic time without sequence-based heuristics. To achieve this low complexity, on par with sequence alignment algorithms, SPARSE features strong sparsification based on structural properties of the RNA ensembles. Following PMcomp, SPARSE gains further speed-up from lightweight energy computation. Although all existing lightweight Sankoff-style methods restrict Sankoff’s original model by disallowing loop deletions and insertions, SPARSE transfers the Sankoff algorithm to the lightweight energy model completely for the first time. Compared with LocARNA, SPARSE achieves similar alignment and better folding quality in significantly less time (speedup: 3.7). At similar run-time, it aligns low sequence identity instances substantially more accurate than RAF, which uses sequence-based heuristics. Availability and implementation: SPARSE is freely available at http://www.bioinf.uni-freiburg.de/Software/SPARSE. Contact: backofen@informatik.uni-freiburg.de Supplementary information: Supplementary data are available at Bioinformatics online.