RNA secondary structure prediction based on free energy and phylogenetic analysis

RNA secondary structure prediction based on free energy and phylogenetic analysis
复制标题

DOI:
10.1006/jmbi.1999.2801
复制
发表时间:
1999-06-18
影响因子:
5.6
通讯作者:
Wilson, C
Wilson, C
中科院分区:
生物学2区
文献类型:
--
作者:
Juan, V;Wilson, C

文献摘要

被引文献

相似文献

我们描述了一种预测RNA二级结构的计算方法,该方法结合了自由能和比较序列分析策略。使用基于同源性的序列比对作为起点,确定相对于特纳能量函数的所有有利配对。多序列比对中的每个潜在配对区域使用将预测自由能和序列共变与优化权重相结合的函数进行评分。高得分区域被排序并顺序并入以定义生长的二级结构。使用一组优化的参数,可以准确预测之前通过广泛的系统发育和实验数据定义的几种测试RNA(包括tRNA、5 S rRNA、SRP RNA、tmRNA和16 S rRNA)的折叠。该算法正确预测了约80%的二级结构。已经测试了一系列参数,以定义准确预测二级结构所需的最小序列信息含量,并评估预测方案中各个术语的重要性。这一分析表明,预测精度最强烈地依赖于协变信息,只有弱的能量项。然而,相对较少的序列证明足以提供准确预测所需的协变信息。二级结构可以通过与少至五个序列的比对来准确地定义,并且预测仅在包含额外序列的情况下适度地改善。(C)北京:科学出版社.
We describe a computational method for the prediction of RNA secondary structure that uses a combination of free energy and comparative sequence analysis strategies. Using a homology-based sequence alignment as a starting point, all favorable pairings with respect to the Turner energy function are identified. Each potentially paired region within a multiple sequence alignment is scored using a function that combines both predicted free energy and sequence covariation with optimized weightings. High scoring regions are ranked and sequentially incorporated to define a growing secondary structure. Using a single set of optimized parameters, it is possible to accurately predict the foldings of several test RNAs defined previously by extensive phylogenetic and experimental data (including tRNA, 5 S rRNA, SRP RNA, tmRNA, and 16 S rRNA). The algorithm correctly predicts approximately 80% of the secondary structure. A range of parameters have been tested to define the minimal sequence information content required to accurately predict secondary structure and to assess the importance of individual terms in the prediction scheme. This analysis indicates that prediction accuracy most strongly depends upon covariational information and only weakly on the energetic terms. However, relatively few sequences prove sufficient to provide the covariational information required for an accurate prediction. Secondary structures can be accurately defined by alignments with as few as five sequences and predictions improve only moderately with the inclusion of additional sequences. (C) 1999 Academic Press.