A max-margin training of RNA secondary structure prediction integrated with the thermodynamic model

A max-margin training of RNA secondary structure prediction integrated with the thermodynamic model
复制标题

DOI:
10.1101/205047
复制
发表时间:
2017-10
期刊:
bioRxiv
影响因子:
--
通讯作者:
Manato Akiyama;Kengo Sato;Y. Sakakibara
Manato Akiyama;Kengo Sato;Y. Sakakibara
中科院分区:
其他
文献类型:
--
作者:
Manato Akiyama;Kengo Sato;Y. Sakakibara

文献摘要

相似文献

动机:预测RNA二级结构的一种流行方法是热力学最近邻模型,该模型找到具有最小自由能(MFE)的热力学最稳定的二级结构。为了进一步改进,开发了一种基于机器学习技术的替代方法。基于机器学习的方法可以使用细粒度模型,该模型包括更丰富的特征表示,具有匹配训练数据的能力。尽管基于机器学习的细粒度模型在预测精度上取得了极高的性能,但有报道称这种模型可能存在过拟合的风险。结果:结合热力学方法和机器学习加权方法,提出了一种新的RNA二级结构预测算法。我们的细粒度模型结合了实验确定的热力学参数和大量针对特征的详细上下文的评分参数,这些参数通过结构化支持向量机和ℓ1正则化进行训练,以避免过拟合。我们的基准测试表明,与现有方法相比,我们的算法达到了最好的预测精度,并且不会出现严重的过拟合。可用性:我们算法的实现可在https://github.com/keio-bioinformatics/mxfold.上获得联系人:Satoken@Bio.keio.ac.jp
Motivation: A popular approach for predicting RNA secondary structure is the thermodynamic nearest neighbor model that finds a thermodynamically most stable secondary structure with the minimum free energy (MFE). For further improvement, an alternative approach that is based on machine learning techniques has been developed. The machine learning based approach can employ a fine-grained model that includes much richer feature representations with the ability to fit the training data. Although a machine learning based fine-grained model achieved extremely high performance in prediction accuracy, a possibility of the risk of overfitting for such model has been reported. Results: In this paper, we propose a novel algorithm for RNA secondary structure prediction that integrates the thermodynamic approach and the machine learning based weighted approach. Ourfine-grained model combines the experimentally determined thermodynamic parameters with a large number of scoring parameters for detailed contexts of features that are trained by the structured support vector machine (SSVM) with the ℓ1 regularization to avoid overfitting. Our benchmark shows that our algorithm achieves the best prediction accuracy compared with existing methods, and heavy overfitting cannot be observed. Availability: The implementation of our algorithm is available at https://github.com/keio-bioinformatics/mxfold. Contact: satoken@bio.keio.ac.jp