Pairwise local structural alignment of RNA sequences with sequence similarity less than 40%

Pairwise local structural alignment of RNA sequences with sequence similarity less than 40%
复制标题

DOI:
10.1093/bioinformatics/bti279
复制
发表时间:
2005-05-01
期刊:
影响因子:
5.8
通讯作者:
Gorodkin, J
Gorodkin, J
中科院分区:
生物学3区
文献类型:
--
作者:
Havgaard, JH;Lyngso, RB;Gorodkin, J

文献摘要

被引文献

相似文献

动机:寻找非编码RNA(ncRNA)基因和结构RNA元件(eleRNA)是当今基因发现的主要挑战,因为这些基因通常在结构上保守而不是在序列上保守。即使可用的方法的数量正在增长,它仍然是感兴趣的成对检测两个基因的序列相似性低,其中的基因是一个较大的基因组region.Results的一部分:在这里,我们提出了这样一种方法,成对的局部比对是基于foldalign和Sankoff算法同时进行多个序列的结构比对。我们包括进行相互扫描的任意长度的两个序列,同时寻找一些最大长度的共同的局部结构图案的能力。这大大降低了算法的复杂性。评分方案包括对应于那些可用于自由能以及类似于RIBOSUM的取代矩阵的结构参数。在数据集上测试新的foldalign实现,其中ncRNA和eleRNA具有< 40%的序列相似性,并且其中ncRNA和eleRNA与周围的基因组序列背景在能量上无法区分。该方法以两种方式进行测试:(1)其仅找到基因之间的共同结构的能力和(2)其在基因组背景中定位ncRNA和eleRNA的能力。在情况(1)中,与Dynalign等方法进行比较是有意义的,性能非常相似,但foldalign要快得多。使用马修斯相关系数,一个家族的结构预测性能通常在0.7左右。在情况(2)中,使用BLAST样命中选择方案,该算法成功定位RNA家族,平均灵敏度为0.8,阳性预测值为0.9。
Motivation: Searching for non-coding RNA (ncRNA) genes and structural RNA elements (eleRNA) are major challenges in gene finding today as these often are conserved in structure rather than in sequence. Even though the number of available methods is growing, it is still of interest to pairwise detect two genes with low sequence similarity, where the genes are part of a larger genomic region.Results: Here we present such an approach for pairwise local alignment which is based on foldalign and the Sankoff algorithm for simultaneous structural alignment of multiple sequences. We include the ability to conduct mutual scans of two sequences of arbitrary length while searching for common local structural motifs of some maximum length. This drastically reduces the complexity of the algorithm. The scoring scheme includes structural parameters corresponding to those available for free energy as well as for substitution matrices similar to RIBOSUM. The new foldalign implementation is tested on a dataset where the ncRNAs and eleRNAs have sequence similarity < 40% and where the ncRNAs and eleRNAs are energetically indistinguishable from the surrounding genomic sequence context. The method is tested in two ways: (1) its ability to find the common structure between the genes only and (2) its ability to locate ncRNAs and eleRNAs in a genomic context. In case (1), it makes sense to compare with methods like Dynalign, and the performances are very similar, but foldalign is substantially faster. The structure prediction performance for a family is typically around 0.7 using Matthews correlation coefficient. In case (2), the algorithm is successful at locating RNA families with an average sensitivity of 0.8 and a positive predictive value of 0.9 using a BLAST-like hit selection scheme.