ExpaRNA-P: simultaneous exact pattern matching and folding of RNAs

ExpaRNA-P: simultaneous exact pattern matching and folding of RNAs
复制标题

DOI:
10.1186/s12859-014-0404-0
复制
发表时间:
2014-12-31
期刊:
影响因子:
3
通讯作者:
Will, Sebastian
Will, Sebastian
中科院分区:
生物学4区
文献类型:
--
作者:
Otto, Christina;Moehl, Mathias;Will, Sebastian

文献摘要

被引文献

相似文献

背景:识别两个 RNA 共有的序列结构基序可以大大加快结构 RNA 的比较速度。现有方法 ExpaRNA 的核心算法针对先验已知的输入结构解决了这个问题。然而,这样的结构却鲜为人知。此外,通过计算来预测它们也无济于事,因为单个序列结构预测非常不可靠。 结果:新颖的算法 ExpaRNA-P 在两个 RNA 的整个玻尔兹曼分布结构集合中计算出完全匹配的序列结构基序;因此,我们同时匹配和折叠RNA,类似于众所周知的RNA“同时对齐和折叠”。虽然这意味着与 ExpaRNA 相比具有更高的灵活性,但 ExpaRNA-P 具有同样非常低的复杂性(时间和空间的二次方),这是通过其新颖的基于结构集成的稀疏化实现的。此外,我们设计了一种通用链接算法来计算 ExpaRNA-P 序列结构基序的兼容子集。我们利用最佳链作为序列结构比对工具 LocARNA 的锚定约束,从而产生了非常快速的 RNA 比对方法 ExpLoc-P。 ExpLoc-P 在多种变体和最先进的方法中进行了基准测试。特别是,我们正式引入并评估问题的严格和宽松变体;后者使得该方法对补偿突变敏感。在一组典型非编码 RNA 的基准测试中,ExpLoc-P 与 LocARNA 具有相似的准确性,但速度快四倍(在两种变体中),同时对于最长的基准序列(大约 400nt),它的速度提高了 30 倍以上。最后,不同的 ExpLoc-P 变体可以根据特定的应用场景定制该方法。 ExpaRNA-P 和 ExpLoc-P 作为 LocARNA 包的一部分进行分发。源代码可在 http://www.bioinf.uni-freiburg.de/Software/ExpaRNA-P 上免费获取。结论:ExpaRNA-P 新颖的基于集合的稀疏化将其复杂性降低到二次时间和空间。因此,ExpaRNA-P 显着加速了序列结构比对,同时保持了比对质量。不同的 ExpaRNA-P 变体支持广泛的应用。
Background: Identifying sequence-structure motifs common to two RNAs can speed up the comparison of structural RNAs substantially. The core algorithm of the existent approach ExpaRNA solves this problem for a priori known input structures. However, such structures are rarely known; moreover, predicting them computationally is no rescue, since single sequence structure prediction is highly unreliable.Results: The novel algorithm ExpaRNA-P computes exactly matching sequence-structure motifs in entire Boltzmann-distributed structure ensembles of two RNAs; thereby we match and fold RNAs simultaneously, analogous to the well-known "simultaneous alignment and folding" of RNAs. While this implies much higher flexibility compared to ExpaRNA, ExpaRNA-P has the same very low complexity (quadratic in time and space), which is enabled by its novel structure ensemble-based sparsification. Furthermore, we devise a generalized chaining algorithm to compute compatible subsets of ExpaRNA-P's sequence-structure motifs. Resulting in the very fast RNA alignment approach ExpLoc-P, we utilize the best chain as anchor constraints for the sequence-structure alignment tool LocARNA. ExpLoc-P is benchmarked in several variants and versus state-of-the-art approaches. In particular, we formally introduce and evaluate strict and relaxed variants of the problem; the latter makes the approach sensitive to compensatory mutations. Across a benchmark set of typical non-coding RNAs, ExpLoc-P has similar accuracy to LocARNA but is four times faster (in both variants), while it achieves a speed-up over 30-fold for the longest benchmark sequences (approximate to 400nt). Finally, different ExpLoc-P variants enable tailoring of the method to specific application scenarios. ExpaRNA-P and ExpLoc-P are distributed as part of the LocARNA package. The source code is freely available at http://www.bioinf.uni-freiburg.de/Software/ExpaRNA-P.Conclusions: ExpaRNA-P's novel ensemble-based sparsification reduces its complexity to quadratic time and space. Thereby, ExpaRNA-P significantly speeds up sequence-structure alignment while maintaining the alignment quality. Different ExpaRNA-P variants support a wide range of applications.