A range of complex probabilistic models for RNA secondary structure prediction that includes the nearest-neighbor model and more

A range of complex probabilistic models for RNA secondary structure prediction that includes the nearest-neighbor model and more
复制标题

DOI:
10.1261/rna.030049.111
复制
发表时间:
2012-02-01
期刊:
RNA
影响因子:
4.5
通讯作者:
Eddy, Sean R.
Eddy, Sean R.
中科院分区:
生物学3区
文献类型:
--
作者:
Rivas, Elena;Lang, Raymond;Eddy, Sean R.

文献摘要

被引文献

相似文献

单序RNA二级结构预测的标准方法使用了最接近的邻居热力学模型,该模型具有数千个实验确定的能量参数。一个有吸引力的替代方法是使用统计方法与结构RNA数据库估计的参数使用统计方法。使用复杂的最近邻居模型(包括Contrafold,Simfold和ContextFold)的判别统计方法报道了良好的结果。尽管概率模型通常更容易训练和使用,但关于生成概率模型(无上下文的语法[SCFGS])的报道很少。为了探索增加复杂性的概率模型,并直接比较概率,热力学和判别方法,我们创建了龙卷风,这是一种计算工具,可以解析广泛的RNA语法架构(包括标准最近的邻居模型,更多)使用可以用概率,能量或任意分数参数化的广义超级准则。通过使用龙卷风,我们发现概率最近的邻居模型的性能与判别方法相当(但不明显优于)。我们发现,复杂的统计模型容易过度拟合RNA结构,并且评估应使用结构非同源训练和测试数据集。过度拟合影响了至少一种已发布的方法(上下文折叠)。改善RNA二级结构预测统计方法的最重要障碍是缺乏当前RNA数据库中良好策划的单序RNA二级结构的多样性。
The standard approach for single-sequence RNA secondary structure prediction uses a nearest-neighbor thermodynamic model with several thousand experimentally determined energy parameters. An attractive alternative is to use statistical approaches with parameters estimated from growing databases of structural RNAs. Good results have been reported for discriminative statistical methods using complex nearest-neighbor models, including CONTRAfold, Simfold, and ContextFold. Little work has been reported on generative probabilistic models (stochastic context-free grammars [SCFGs]) of comparable complexity, although probabilistic models are generally easier to train and to use. To explore a range of probabilistic models of increasing complexity, and to directly compare probabilistic, thermodynamic, and discriminative approaches, we created TORNADO, a computational tool that can parse a wide spectrum of RNA grammar architectures (including the standard nearest-neighbor model and more) using a generalized super-grammar that can be parameterized with probabilities, energies, or arbitrary scores. By using TORNADO, we find that probabilistic nearest-neighbor models perform comparably to (but not significantly better than) discriminative methods. We find that complex statistical models are prone to overfitting RNA structure and that evaluations should use structurally nonhomologous training and test data sets. Overfitting has affected at least one published method (ContextFold). The most important barrier to improving statistical approaches for RNA secondary structure prediction is the lack of diversity of well-curated single-sequence RNA secondary structures in current RNA databases.