Efficient parameter estimation for RNA secondary structure prediction

Efficient parameter estimation for RNA secondary structure prediction
复制标题

DOI:
10.1093/bioinformatics/btm223
复制
发表时间:
2007-07-01
期刊:
影响因子:
5.8
通讯作者:
Murphy, Kevin P.
Murphy, Kevin P.
中科院分区:
生物学3区
文献类型:
--
作者:
Andronescu, Mirela;Condon, Anne;Murphy, Kevin P.

文献摘要

被引文献

相似文献

动机:根据碱基序列准确预测 RNA 二级结构是一个尚未解决的计算挑战。自由能最小化所做的预测的准确性受到基础自由能模型中能量参数质量的限制。使用最广泛的模型 Turner99 模型具有数百个参数,因此强大的参数估计方案应该有效地处理具有数千个结构的大型数据集。此外,除了结构数据之外,还应该使用可用的实验自由能数据来训练估计方案。结果:在这项工作中,我们提出了约束生成(CG),这是第一个 RNA 自由能参数估计的计算方法,可以在大量结构数据和热力学数据上进行有效训练。我们的 CG 方法采用了一种新颖的迭代方案,首先计算能量值作为约束优化问题的解。然后利用新计算出的能量参数来更新优化函数的约束,以便在下一次迭代中更好地优化能量参数。使用我们对生物健全数据的方法,我们获得了 Turner99 能量模型的修正参数。我们表明,通过使用我们的新参数,我们在预测准确性方面比当前最先进的方法获得了显着提高。
Motivation: Accurate prediction of RNA secondary structure from the base sequence is an unsolved computational challenge. The accuracy of predictions made by free energy minimization is limited by the quality of the energy parameters in the underlying free energy model. The most widely used model, the Turner99 model, has hundreds of parameters, and so a robust parameter estimation scheme should efficiently handle large data sets with thousands of structures. Moreover, the estimation scheme should also be trained using available experimental free energy data in addition to structural data.Results: In this work, we present constraint generation (CG), the first computational approach to RNA free energy parameter estimation that can be efficiently trained on large sets of structural as well as thermodynamic data. Our CG approach employs a novel iterative scheme, whereby the energy values are first computed as the solution to a constrained optimization problem. Then the newly computed energy parameters are used to update the constraints on the optimization function, so as to better optimize the energy parameters in the next iteration. Using our method on biologically sound data, we obtain revised parameters for the Turner99 energy model. We show that by using our new parameters, we obtain significant improvements in prediction accuracy over current state of-the-art methods.