Data-directed RNA secondary structure prediction using probabilistic modeling.

Data-directed RNA secondary structure prediction using probabilistic modeling.
复制标题

DOI:
10.1261/rna.055756.115
复制
发表时间:
2016-08
期刊:
RNA (New York, N.Y.)
影响因子:
--
通讯作者:
Aviran S
Aviran S
中科院分区:
其他
文献类型:
--
作者:
Deng F;Ledda M;Vaziri S;Aviran S

文献摘要

被引文献

相似文献

结构决定了许多RNA的功能,但二级RNA结构分析要么是劳动密集型和昂贵的,要么依赖于经常不准确的计算预测。通过将结构探测数据集成到预测算法中来减轻这些限制。然而,现有算法针对特定类型的探测数据进行了优化。最近,新的化学与测序的进步相结合,以前所未有的规模和灵敏度促进了结构探测。这些新技术和预期的大量数据突出了对易于适应更复杂和多样化输入源的算法的需求。我们实施和研究了最近概述的RNA二级结构预测的概率框架,并将其扩展以适应结构信息的进一步细化。该框架利用每个所考虑的结构上下文的伪能量项的直接基于似然的计算,并且可以容易地适应不同的数据类型和复杂的数据依赖性。我们使用真实的数据结合模拟来评估几种实现的性能,并表明适当的结构上下文集成可以导致改进。我们的测试还揭示了真实的数据和模拟之间的差异,我们表明可以通过精细建模来缓解。然后,我们提出了统计预处理方法来标准化数据解释和集成到这样一个通用的框架。我们进一步系统地量化了数据子集的信息内容,表明高反应性是SHAPE导向预测的主要驱动因素,更好地理解信息量较少的反应性是进一步改进的关键。最后,我们提供的证据,我们的框架使用模拟探针模拟的自适应能力。
Structure dictates the function of many RNAs, but secondary RNA structure analysis is either labor intensive and costly or relies on computational predictions that are often inaccurate. These limitations are alleviated by integration of structure probing data into prediction algorithms. However, existing algorithms are optimized for a specific type of probing data. Recently, new chemistries combined with advances in sequencing have facilitated structure probing at unprecedented scale and sensitivity. These novel technologies and anticipated wealth of data highlight a need for algorithms that readily accommodate more complex and diverse input sources. We implemented and investigated a recently outlined probabilistic framework for RNA secondary structure prediction and extended it to accommodate further refinement of structural information. This framework utilizes direct likelihood-based calculations of pseudo-energy terms per considered structural context and can readily accommodate diverse data types and complex data dependencies. We use real data in conjunction with simulations to evaluate performances of several implementations and to show that proper integration of structural contexts can lead to improvements. Our tests also reveal discrepancies between real data and simulations, which we show can be alleviated by refined modeling. We then propose statistical preprocessing approaches to standardize data interpretation and integration into such a generic framework. We further systematically quantify the information content of data subsets, demonstrating that high reactivities are major drivers of SHAPE-directed predictions and that better understanding of less informative reactivities is key to further improvements. Finally, we provide evidence for the adaptive capability of our framework using mock probe simulations.