CONTRAfold: RNA secondary structure prediction without physics-based models

CONTRAfold: RNA secondary structure prediction without physics-based models
复制标题

DOI:
10.1093/bioinformatics/btl246
复制
发表时间:
2006-07-01
期刊:
影响因子:
5.8
通讯作者:
Batzoglou, Serafim
Batzoglou, Serafim
中科院分区:
生物学3区
文献类型:
--
作者:
Do, Chuong B.;Woods, Daniel A.;Batzoglou, Serafim

文献摘要

被引文献

相似文献

动机:几十年来,自由能最小化方法一直是单序列RNA二级结构预测的主导策略。最近,随机上下文无关文法(SCFG)已经成为一种替代的概率方法来建模RNA结构。与基于物理的方法不同,它依赖于数千个实验测量的热力学参数,SCFG使用全自动统计学习算法来获得模型参数。然而,尽管有这个优势,概率方法并没有取代自由能最小化方法作为二级结构预测的首选工具,因为目前最好的SCFG的准确性还没有与最好的基于物理的模型相匹配。本文提出了一种基于条件对数线性模型(CLLM)的二级结构预测新方法--一类灵活的概率模型,通过使用区分训练和功能丰富的评分来概括SCFG。在一系列的交叉验证实验中,我们表明,基于语法的二级结构预测方法制定CLLM一贯优于他们的SCFG类似物。此外,CLLM结合了典型热力学模型中的大多数特征,实现了迄今为止最高的单序列预测精度,优于目前可用的概率和基于物理的技术。因此,我们的结果关闭概率和热力学模型之间的差距,表明统计学习程序提供了一个有效的替代经验测量的RNA二级结构预测的热力学参数。
Motivation: For several decades, free energy minimization methods have been the dominant strategy for single sequence RNA secondary structure prediction. More recently, stochastic context-free grammars (SCFGs) have emerged as an alternative probabilistic methodology for modeling RNA structure. Unlike physics-based methods, which rely on thousands of experimentally-measured thermodynamic parameters, SCFGs use fully-automated statistical learning algorithms to derive model parameters. Despite this advantage, however, probabilistic methods have not replaced free energy minimization methods as the tool of choice for secondary structure prediction, as the accuracies of the best current SCFGs have yet to match those of the best physics-based models.Results: In this paper, we present CONTRAfold, a novel secondary structure prediction method based on conditional log-linear models (CLLMs), a flexible class of probabilistic models which generalize upon SCFGs by using discriminative training and feature-rich scoring. In a series of cross-validation experiments, we show that grammar-based secondary structure prediction methods formulated as CLLMs consistently outperform their SCFG analogs. Furthermore, CONTRAfold, a CLLM incorporating most of the features found in typical thermodynamic models, achieves the highest single sequence prediction accuracies to date, outperforming currently available probabilistic and physics-based techniques. Our result thus closes the gap between probabilistic and thermodynamic models, demonstrating that statistical learning procedures provide an effective alternative to empirical measurement of thermodynamic parameters for RNA secondary structure prediction.