Dynalign: An algorithm for finding the secondary structure common to two RNA sequences

Dynalign: An algorithm for finding the secondary structure common to two RNA sequences
复制标题

DOI:
10.1006/jmbi.2001.5351
复制
发表时间:
2002-03-22
影响因子:
5.6
通讯作者:
Turner, DH
Turner, DH
中科院分区:
生物学2区
文献类型:
--
作者:
Mathews, DH;Turner, DH

文献摘要

被引文献

相似文献

随着基因组序列数据库规模的快速增长,RNA的计算分析在揭示结构-功能关系和潜在药物靶点方面将变得越来越重要。对于已知二级结构的大型数据库,单个序列的RNA二级结构预测平均准确率为73%。这种水平的准确性提供了一个很好的起点,无论是通过比较序列分析或实验研究的解释确定二级结构。Dynalign是一种新的计算机算法,它通过结合自由能最小化和比较序列分析来找到两个序列共同的低自由能结构,而不需要任何序列同一性,从而提高了结构预测的准确性。它使用了Sankoff建议的动态规划结构。然而,Dynalign限制了两个序列中比对核苷酸之间允许的最大距离M。这使得计算易于处理,因为复杂性被简化为O((MN 3)-N-3),其中N是较短序列的长度。Dynalign的准确性用13个tRNA、7个5 S rRNA和2个R2 3' UTR序列的组进行测试。平均而言,Dynalign预测了tRNA中86.1%的已知碱基对,而单独使用自由能最小化则为59.7%。对于5SrRNA,平均准确率从47.8%提高到86.4%。来自Drosophila takahashii的R2 3' UTR的二级结构通过标准自由能最小化预测不佳。然而,使用Dynalign,与来自黑腹果蝇的序列串联预测的结构几乎与通过比较序列分析确定的结构匹配。(C)2002 Elsevier Science Ltd.
With the rapid increase in the size of the genome sequence database, computational analysis of RNA will become increasingly important in revealing structure-function relationships and potential drug targets. RNA secondary structure prediction for a single sequence is 73% accurate on average for a large database of known secondary structures. This level of accuracy provides a good starting point for determining a secondary structure either by comparative sequence analysis or by the interpretation of experimental studies. Dynalign is a new computer algorithm that improves the accuracy of structure prediction by combining free energy minimization and comparative sequence analysis to find a low free energy structure common to two sequences without requiring any sequence identity. It uses a dynamic programming construct suggested by Sankoff. Dynalign, however, restricts the maximum distance, M, allowed between aligned nucleotides in the two sequences. This makes the calculation tractable because the complexity is simplified to O((MN3)-N-3), where N is the length of the shorter sequence.The accuracy of Dynalign was tested with sets of 13 tRNAs, seven 5 S rRNAs, and two R2 3' UTR sequences. On average, Dynalign predicted 86.1% of known base-pairs in the tRNAs, as compared to 59.7% for free energy minimization alone. For the 5 S rRNAs, the average accuracy improves from 47.8% to 86.4%. The secondary structure of the R2 3' UTR from Drosophila takahashii is poorly predicted by standard free energy minimization. With Dynalign, however, the structure predicted in tandem with the sequence from Drosophila melanogaster nearly matches the structure determined by comparative sequence analysis. (C) 2002 Elsevier Science Ltd.