Efficient pairwise RNA structure prediction using probabilistic alignment constraints in Dynalign.

Efficient pairwise RNA structure prediction using probabilistic alignment constraints in Dynalign.
复制标题

DOI:
10.1186/1471-2105-8-130
复制
发表时间:
2007-04-19
期刊:
影响因子:
3
通讯作者:
Mathews DH
Mathews DH
中科院分区:
生物学4区
文献类型:
--
作者:
Harmanci AO;Sharma G;Mathews DH

文献摘要

参考文献

被引文献

相似文献

两个RNA序列的联合比对和二级结构预测可以显着提高结构预测的准确性。然而,解决这个问题的方法被迫采用约束,通过限制允许的对齐和/或结构(即折叠)来减少计算。在本文中,提出了一种新的方法,旨在建立基于核苷酸比对和插入后验概率的比对约束。使用隐马尔可夫模型,计算两个序列中核苷酸位置的所有可能配对的比对和插入的后验概率。这些比对和插入后验概率相加组合以获得核苷酸位置对的重合概率。通过对重合概率进行阈值处理来获得合适的对齐约束。该约束与 Dynalign 集成,Dynalign 是一种用于关节对齐和二级结构预测的自由能最小化算法。由此产生的方法以之前版本的 Dynalign 和其他成对 RNA 结构预测程序为基准。所提出的技术消除了 Dynalign 中的手动参数选择,与 Dynalign 中的先前约束相比,显着节省了计算时间,同时在结构预测精度方面提供了小幅改进。内存中也实现了节省。在平均序列长度约为 120 个核苷酸的 5S RNA 数据集上进行的实验中,该方法将计算量减少了 2 倍。与其他成对 RNA 结构预测程序相比,该方法表现良好:平均而言,产生更好的准确性,并且需要的计算资源显着减少。可以利用概率分析以原则性方式自动确定成对 RNA 结构预测方法的比对约束。这些约束可以减少这些方法的计算和内存需求,同时保持或提高其结构预测的准确性。这将这些方法的实际范围扩展到更长的序列。修订后的 Dynaalign 代码可免费下载。
Joint alignment and secondary structure prediction of two RNA sequences can significantly improve the accuracy of the structural predictions. Methods addressing this problem, however, are forced to employ constraints that reduce computation by restricting the alignments and/or structures (i.e. folds) that are permissible. In this paper, a new methodology is presented for the purpose of establishing alignment constraints based on nucleotide alignment and insertion posterior probabilities. Using a hidden Markov model, posterior probabilities of alignment and insertion are computed for all possible pairings of nucleotide positions from the two sequences. These alignment and insertion posterior probabilities are additively combined to obtain probabilities of co-incidence for nucleotide position pairs. A suitable alignment constraint is obtained by thresholding the co-incidence probabilities. The constraint is integrated with Dynalign, a free energy minimization algorithm for joint alignment and secondary structure prediction. The resulting method is benchmarked against the previous version of Dynalign and against other programs for pairwise RNA structure prediction. The proposed technique eliminates manual parameter selection in Dynalign and provides significant computational time savings in comparison to prior constraints in Dynalign while simultaneously providing a small improvement in the structural prediction accuracy. Savings are also realized in memory. In experiments over a 5S RNA dataset with average sequence length of approximately 120 nucleotides, the method reduces computation by a factor of 2. The method performs favorably in comparison to other programs for pairwise RNA structure prediction: yielding better accuracy, on average, and requiring significantly lesser computational resources. Probabilistic analysis can be utilized in order to automate the determination of alignment constraints for pairwise RNA structure prediction methods in a principled fashion. These constraints can reduce the computational and memory requirements of these methods while maintaining or improving their accuracy of structural prediction. This extends the practical reach of these methods to longer length sequences. The revised Dynalign code is freely available for download.
DOI: 10.1504/ijbra.2005.007581
发表时间: 2005-01-01
影响因子: --
作者:
Masoumi, Beeta;Turcotte, Marcel
通讯作者: Turcotte, Marcel
DOI: 10.1093/nar/28.4.991
发表时间: 2000-02-15
影响因子: 14.9
作者:
Chen, JH;Le, SY;Maizel, JV
通讯作者: Maizel, JV
DOI: 10.1093/nar/gkg006
发表时间: 2003-01-01
影响因子: 14.9
作者:
Griffiths-Jones, S;Bateman, A;Eddy, SR
通讯作者: Eddy, SR
DOI: 10.1016/s0959-440x(02)00339-1
发表时间: 2002-06-01
影响因子: 6.8
作者:
Gutell, RR;Lee, JC;Cannone, JJ
通讯作者: Cannone, JJ
DOI: 10.1093/bioinformatics/bti349
发表时间: 2005-05-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Mathews, DH
通讯作者: Mathews, DH