Improved prediction of RNA secondary structure by integrating the free energy model with restraints derived from experimental probing data.

Improved prediction of RNA secondary structure by integrating the free energy model with restraints derived from experimental probing data.
复制标题

通过将自由能模型与来自实验探测数据的约束相结合,改进了 RNA 二级结构的预测

DOI:
10.1093/nar/gkv706
复制
发表时间:
2015-09-03
影响因子:
14.9
通讯作者:
Lu ZJ
Lu ZJ
中科院分区:
生物学2区
文献类型:
--
作者:
Wu Y;Shi B;Ding X;Liu T;Hu X;Yip KY;Yang ZR;Mathews DH;Lu ZJ

文献摘要

被引文献

相似文献

最近,出现了几种基于高通量测序的探测RNA结构的实验技术。然而,大多数包含探测数据的二级结构预测工具是针对特定类型的实验设计和优化的。例如,RNAstructure-Fold针对SHAPE数据进行优化,而SeqFold针对PARS数据进行优化。在这里,我们报告了一种新的RNA二级结构预测方法,限制MaxExpect(RME),它可以结合多种类型的实验探测数据,是基于自由能模型和MEA(最大化预期精度)算法。我们首先证明,RME大大提高了二级结构预测与完美的限制(已知结构的碱基对信息)。接下来,我们收集了来自不同实验(例如SHAPE,PARS和DMS-seq)的结构探测数据,并将其转换为一组统一的配对概率与后验概率模型。通过使用概率分数作为RME中的约束,我们将其二级结构预测性能与其他两种众所周知的工具RNAstructure-Fold(基于自由能最小化算法)和SeqFold(基于采样算法)进行了比较。对于SHAPE数据,RME和RNAstructure-Fold的表现优于SeqFold,因为它们显著改变了具有实验约束的能量模型。对于具有较低探测效率的高通量数据(例如PARS和DMS-seq),测试工具的二级结构预测性能是相当的,仅测试RNA的一部分具有性能改进。然而,当去除三级结构和蛋白质相互作用的影响时,通过结合体内DMS-seq数据,RME在DMS可接近区域中显示出最高的预测准确性。
Recently, several experimental techniques have emerged for probing RNA structures based on high-throughput sequencing. However, most secondary structure prediction tools that incorporate probing data are designed and optimized for particular types of experiments. For example, RNAstructure-Fold is optimized for SHAPE data, while SeqFold is optimized for PARS data. Here, we report a new RNA secondary structure prediction method, restrained MaxExpect (RME), which can incorporate multiple types of experimental probing data and is based on a free energy model and an MEA (maximizing expected accuracy) algorithm. We first demonstrated that RME substantially improved secondary structure prediction with perfect restraints (base pair information of known structures). Next, we collected structure-probing data from diverse experiments (e.g. SHAPE, PARS and DMS-seq) and transformed them into a unified set of pairing probabilities with a posterior probabilistic model. By using the probability scores as restraints in RME, we compared its secondary structure prediction performance with two other well-known tools, RNAstructure-Fold (based on a free energy minimization algorithm) and SeqFold (based on a sampling algorithm). For SHAPE data, RME and RNAstructure-Fold performed better than SeqFold, because they markedly altered the energy model with the experimental restraints. For high-throughput data (e.g. PARS and DMS-seq) with lower probing efficiency, the secondary structure prediction performances of the tested tools were comparable, with performance improvements for only a portion of the tested RNAs. However, when the effects of tertiary structure and protein interactions were removed, RME showed the highest prediction accuracy in the DMS-accessible regions by incorporating in vivo DMS-seq data.