Prediction and statistics of pseudoknots in RNA structures using exactly clustered stochastic simulations

Prediction and statistics of pseudoknots in RNA structures using exactly clustered stochastic simulations
复制标题

DOI:
10.1073/pnas.2536430100
复制
发表时间:
2003-12-23
影响因子:
11.1
通讯作者:
Isambert, H
Isambert, H
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Xayaphoummine, A;Bucher, T;Isambert, H

文献摘要

被引文献

相似文献

从头算RNA二级结构预测长期以来一直拒绝考虑环内部的螺旋,即所谓的假结,尽管它们在结构上很重要。在这里,我们报告说,许多假结点可以通过长时间尺度的RNA折叠模拟来预测,该模拟遵循单个RNA螺旋的随机关闭和打开。这些随机模拟的数值有效性依赖于一种theta(n(2))聚类算法,该算法在一组连续更新的n个参考结构上计算时间平均值。应用这种精确的随机聚类方法,我们通常可以获得高达400个碱基的RNA序列的5到100倍的模拟加速,而对于短的、多稳定的分子(小于或等于150个碱基),有效的加速可以高达10(5)倍。我们对随机和自然的RNA序列进行了广泛的折叠统计,发现伪结点在RNA结构中分布不均匀,占富含G+C的RNA序列碱基对的30%(在线RNA折叠动力学服务器,包括伪结点:http://kinefold.u-strasbg.fr)。
Ab initio RNA secondary structure predictions have long dismissed helices interior to loops, so-called pseudoknots, despite their structural importance. Here we report that many pseudoknots can be predicted through long-time-scale RNA-folding simulations, which follow the stochastic closing and opening of individual RNA helices. The numerical efficacy of these stochastic simulations relies on an theta(n(2)) clustering algorithm that computes time averages over a continuously updated set of n reference structures. Applying this exact stochastic clustering approach, we typically obtain a 5- to 100-fold simulation speed-up for RNA sequences up to 400 bases, while the effective acceleration can be as high as 10(5)-fold for short, multistable molecules (less than or equal to150 bases). We performed extensive folding statistics on random and natural RNA sequences and found that pseudoknots are distributed unevenly among RNA structures and account for up to 30% of base pairs in G+C-rich RNA sequences (online RNA-folding kinetics server including pseudoknots: http:// kinefold.u-strasbg.fr).