An Iterated loop matching approach to the prediction of RNA secondary structures with pseudoknots

An Iterated loop matching approach to the prediction of RNA secondary structures with pseudoknots
复制标题

DOI:
10.1093/bioinformatics/btg373
复制
发表时间:
2004-01-01
期刊:
影响因子:
5.8
通讯作者:
Zhang, WX
Zhang, WX
中科院分区:
生物学3区
文献类型:
--
作者:
Ruan, JH;Stormo, GD;Zhang, WX

文献摘要

被引文献

相似文献

动机:由于其难以建模,伪结通常被排除在RNA二级结构的预测之外。虽然,有几个动态规划算法存在的预测伪结使用热力学方法,他们既不可靠,也没有效率。另一方面,比较方法更可靠,但往往是以临时方式进行的,需要专家干预。最大加权匹配,一个算法的pseudoknot预测与比较分析,遭受低预测精度在许多cases.Results:在这里,我们提出了一种算法,迭代循环匹配,可靠和有效地预测RNA二级结构,包括pseudoknot。该方法可以利用热力学或比较信息或两者,因此能够预测比对序列和单个序列的假结。我们已经测试了一些RNA家族的算法。使用8-12个同源序列,该算法正确识别了短序列的90%以上的碱基对和80%的总体碱基对。它正确地预测几乎所有的伪结,并产生很少的假碱基对序列没有伪结。比较表明,我们的算法是更敏感和更具体的最大加权匹配方法。此外,我们的算法具有较高的预测精度的个别序列,与PKNOTS算法相比,而使用更少的计算资源。
Motivation: Pseudoknots have generally been excluded from the prediction of RNA secondary structures due to its difficulty in modeling. Although, several dynamic programming algorithms exist for the prediction of pseudoknots using thermodynamic approaches, they are neither reliable nor efficient. On the other hand, comparative methods are more reliable, but are often done in an ad hoc manner and require expert intervention. Maximum weighted matching, an algorithm for pseudoknot prediction with comparative analysis, suffers from low-prediction accuracy in many cases.Results: Here we present an algorithm, iterated loop matching, for reliably and efficiently predicting RNA secondary structures including pseudoknots. The method can utilize either thermodynamic or comparative information or both, thus is able to predict pseudoknots for both aligned and individual sequences. We have tested the algorithm on a number of RNA families. Using 8-12 homologous sequences, the algorithm correctly identifies more than 90% of base-pairs for short sequences and 80% overall. It correctly predicts nearly all pseudoknots and produces very few spurious base-pairs for sequences without pseudoknots. Comparisons show that our algorithm is both more sensitive and more specific than the maximum weighted matching method. In addition, our algorithm has high-prediction accuracy on individual sequences, comparable with the PKNOTS algorithm, while using much less computational resources.