Rich Parameterization Improves RNA Structure Prediction

Rich Parameterization Improves RNA Structure Prediction
复制标题

DOI:
10.1089/cmb.2011.0184
复制
发表时间:
2011-11-01
影响因子:
1.7
通讯作者:
Ziv-Ukelson, Michal
Ziv-Ukelson, Michal
中科院分区:
生物学4区
文献类型:
--
作者:
Zakov, Shay;Goldberg, Yoav;Ziv-Ukelson, Michal

文献摘要

被引文献

相似文献

目前预测RNA结构的方法从基于物理的方法到机器学习(ML)技术,这些方法依赖于数千个实验测量的热力学参数。虽然参数估计的方法正在成功地转向基于ML的方法,但模型的参数化到目前为止仍然相当稳定。我们研究了增加RNA折叠预测模型利用的信息量对提高其预测质量的潜在贡献。这是通过提出新的模型来实现的,该模型通过检查更多类型的结构元素来改进以前的模型,并为这些元素提供更大的顺序上下文。我们提出的细粒度模型变得实用,这要归功于大训练集的可用性、机器学习的进步以及最近对RNA折叠算法的加速。结果表明,更详细的模型的应用确实提高了预测质量,而折叠算法的相应运行时间保持较快。这个实验的另一个重要结果是一个新的RNA折叠预测模型(加上一个免费的实现),它的预测质量比以前的模型高得多。这个最终的模型有大约7万个自由参数,比以前的模型多了几个数量级。在相同的综合数据集上进行训练和测试,我们的模型在正确预测的碱基对上根据F-1度量获得了84%的分数(即16%的错误率),而之前的最好报告分数为70%(即30%的错误率)。也就是说,新模型的误差减少了约50%。训练有素的模型和源代码可在www.cs.bgu.ac.il/上找到,类似于Negevcb/ConextFold。
Current approaches to RNA structure prediction range from physics-based methods, which rely on thousands of experimentally measured thermodynamic parameters, to machine-learning (ML) techniques. While the methods for parameter estimation are successfully shifting toward ML-based approaches, the model parameterizations so far remained fairly constant. We study the potential contribution of increasing the amount of information utilized by RNA folding prediction models to the improvement of their prediction quality. This is achieved by proposing novel models, which refine previous ones by examining more types of structural elements, and larger sequential contexts for these elements. Our proposed fine-grained models are made practical thanks to the availability of large training sets, advances in machine-learning, and recent accelerations to RNA folding algorithms. We show that the application of more detailed models indeed improves prediction quality, while the corresponding running time of the folding algorithm remains fast. An additional important outcome of this experiment is a new RNA folding prediction model (coupled with a freely available implementation), which results in a significantly higher prediction quality than that of previous models. This final model has about 70,000 free parameters, several orders of magnitude more than previous models. Being trained and tested over the same comprehensive data sets, our model achieves a score of 84% according to the F-1-measure over correctly-predicted base-pairs (i.e., 16% error rate), compared to the previously best reported score of 70% (i.e., 30% error rate). That is, the new model yields an error reduction of about 50%. Trained models and source code are available at www.cs.bgu.ac.il/similar to negevcb/contextfold.