Multithreaded comparative RNA secondary structure prediction using stochastic context-free grammars.

Multithreaded comparative RNA secondary structure prediction using stochastic context-free grammars.
复制标题

DOI:
10.1186/1471-2105-12-103
复制
发表时间:
2011-04-18
期刊:
影响因子:
3
通讯作者:
Andersen ES
Andersen ES
中科院分区:
生物学4区
文献类型:
--
作者:
Sükösd Z;Knudsen B;Vaerum M;Kjems J;Andersen ES

文献摘要

参考文献

被引文献

相似文献

大RNA结构的预测仍然是生物信息学中的一个特殊挑战,这是由于计算复杂性和最先进算法的低准确度。pfold模型将随机上下文无关语法与系统发育分析相结合,以获得高精度的预测,但算法的时间复杂性和下溢错误阻止了其用于长比对。在这里,我们提出了PPfold,pfold的多线程版本,这是能够预测大RNA比对的结构,准确地在实际的时间尺度。我们已经在PPfold中分发了系统发育计算和内部-外部算法,从而大大减少了多核机器上的运行时间。我们通过实现扩展指数数据类型解决了pfold的浮点下溢问题,使PPfold能够用于大规模RNA结构预测。我们还改进了用户界面和可移植性:除了程序的独立可执行文件和Java源代码外,PPfold还可以作为CLC工作台的免费插件。我们使用BRaliBase I测试评估了PPfold的准确性,并通过在8核机器上在65分钟内预测24个完整HIV-1基因组的二级结构并识别预测中的几个已知结构元素来证明其实际用途。PPfold是迄今为止第一个并行比较RNA结构预测算法。基于pfold模型,PPfold能够快速、高质量地预测大型RNA二级结构,例如RNA病毒的基因组或长基因组转录物。该算法的并行化中所使用的技术可能对其他生物信息学算法具有普遍适用性。
The prediction of the structure of large RNAs remains a particular challenge in bioinformatics, due to the computational complexity and low levels of accuracy of state-of-the-art algorithms. The pfold model couples a stochastic context-free grammar to phylogenetic analysis for a high accuracy in predictions, but the time complexity of the algorithm and underflow errors have prevented its use for long alignments. Here we present PPfold, a multithreaded version of pfold, which is capable of predicting the structure of large RNA alignments accurately on practical timescales. We have distributed both the phylogenetic calculations and the inside-outside algorithm in PPfold, resulting in a significant reduction of runtime on multicore machines. We have addressed the floating-point underflow problems of pfold by implementing an extended-exponent datatype, enabling PPfold to be used for large-scale RNA structure predictions. We have also improved the user interface and portability: alongside standalone executable and Java source code of the program, PPfold is also available as a free plugin to the CLC Workbenches. We have evaluated the accuracy of PPfold using BRaliBase I tests, and demonstrated its practical use by predicting the secondary structure of an alignment of 24 complete HIV-1 genomes in 65 minutes on an 8-core machine and identifying several known structural elements in the prediction. PPfold is the first parallelized comparative RNA structure prediction algorithm to date. Based on the pfold model, PPfold is capable of fast, high-quality predictions of large RNA secondary structures, such as the genomes of RNA viruses or long genomic transcripts. The techniques used in the parallelization of this algorithm may be of general applicability to other bioinformatics algorithms.
DOI: 10.1007/bf00818163
发表时间: 1994-02-01
影响因子: 1.8
作者:
HOFACKER, IL;FONTANA, W;SCHUSTER, P
通讯作者: SCHUSTER, P
DOI: 10.1093/nar/gkg614
发表时间: 2003-07-01
影响因子: 14.9
作者:
Knudsen, B;Hein, J
通讯作者: Hein, J
DOI: 10.1093/bioinformatics/15.6.446
发表时间: 1999-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Knudsen, B;Hein, J
通讯作者: Hein, J
DOI: 10.1089/10665270050081441
发表时间: 2000-02-01
影响因子: 1.7
作者:
Fekete, M;Hofacker, IL;Stadler, PF
通讯作者: Stadler, PF
DOI: 10.1186/1471-2105-5-140
发表时间: 2004-09-30
期刊: BMC bioinformatics
影响因子: 3
作者:
Gardner PP;Giegerich R
通讯作者: Giegerich R