Searching for evolutionary distant RNA homologs within genomic sequences using partition function posterior probabilities.

Searching for evolutionary distant RNA homologs within genomic sequences using partition function posterior probabilities.
复制标题

DOI:
10.1186/1471-2105-9-61
复制
发表时间:
2008-01-28
期刊:
影响因子:
3
通讯作者:
Livesay DR
Livesay DR
中科院分区:
生物学4区
文献类型:
--
作者:
Roshan U;Chikkagoudar S;Livesay DR

文献摘要

参考文献

被引文献

相似文献

当成对序列同一性低或存在不可替代的侧翼残基时,基因组区段内的RNA同源物的鉴定是困难的。在这两种情况下,结构-序列或谱/家族-序列比对程序变得难以应用,因为不可靠的RNA结构或家族比对。因此,经常使用局部序列-序列比对程序。我们最近已经证明,使用分区函数匹配概率(在Probalign中实现)的最大预期准确度比对在异质长度蛋白质序列数据集上明显优于当代方法,从而表明对局部比对的亲和力。我们从RFAM家族中创建了平均成对序列同一性高达60%的成对RNA基因组比对基准。每个数据集包含与嵌入在至少5 K个核苷酸长的基因组序列中的靶RNA(相同家族的)比对的查询RNA。为了模拟当ncRNA的确切末端未知时的常见条件,每个查询RNA具有大小为50、100和150个核苷酸的5'和3'基因组侧翼。随后,我们比较的错误的Probalign程序(调整本地对齐),常用的本地对齐程序HMMER,SHELL,和BLAST,和流行的ClustalW程序与零末端缺口处罚。每个程序的参数都在基准测试的一个小子集上进行了优化。Probalign在整个基准测试中具有最高的准确性。它在22个家庭中的5个家庭中的准确率超过了Scrum(下一个最好的方法)。在限制于最大30%序列同一性的数据集上,Probalign的总体中位误差为71.2%,而Scrum为83.4%(P值< 0.05)。此外,在这些数据集上,Probalign在五个家庭中至少领先Sunday 10%; Sunday在十四个家庭中的两个家庭中领先Probalign相同的幅度。我们还证明了Probalign平均后验概率,相比于归一化的SZ-score,是比对质量的更好的估计。所有数据集和软件均可在线获取。我们证明,第一次,分区函数匹配概率用于预期的准确性比对,如在Probalign,提供了统计学上显着的改善,目前的方法,用于确定在较大的基因组片段中的远亲RNA序列。
Identification of RNA homologs within genomic stretches is difficult when pairwise sequence identity is low or unalignable flanking residues are present. In both cases structure-sequence or profile/family-sequence alignment programs become difficult to apply because of unreliable RNA structures or family alignments. As such, local sequence-sequence alignment programs are frequently used instead. We have recently demonstrated that maximal expected accuracy alignments using partition function match probabilities (implemented in Probalign) are significantly better than contemporary methods on heterogeneous length protein sequence datasets, thus suggesting an affinity for local alignment. We create a pairwise RNA-genome alignment benchmark from RFAM families with average pairwise sequence identity up to 60%. Each dataset contains a query RNA aligned to a target RNA (of the same family) embedded in a genomic sequence at least 5K nucleotides long. To simulate common conditions when exact ends of an ncRNA are unknown, each query RNA has 5' and 3' genomic flanks of size 50, 100, and 150 nucleotides. We subsequently compare the error of the Probalign program (adjusted for local alignment) to the commonly used local alignment programs HMMER, SSEARCH, and BLAST, and the popular ClustalW program with zero end-gap penalties. Parameters were optimized for each program on a small subset of the benchmark. Probalign has overall highest accuracies on the full benchmark. It leads by 10% accuracy over SSEARCH (the next best method) on 5 out of 22 families. On datasets restricted to maximum of 30% sequence identity, Probalign's overall median error is 71.2% vs. 83.4% for SSEARCH (P-value < 0.05). Furthermore, on these datasets Probalign leads SSEARCH by at least 10% on five families; SSEARCH leads Probalign by the same margin on two of the fourteen families. We also demonstrate that the Probalign mean posterior probability, compared to the normalized SSEARCH Z-score, is a better discriminator of alignment quality. All datasets and software are available online. We demonstrate, for the first time, that partition function match probabilities used for expected accuracy alignment, as done in Probalign, provide statistically significant improvement over current approaches for identifying distantly related RNA sequences in larger genomic segments.
DOI: 10.1016/s0097-8485(96)80004-0
发表时间: 1996-03-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gribskov, M;Robinson, NL
通讯作者: Robinson, NL
DOI: 10.1093/protein/8.10.999
发表时间: 1995-10-01
期刊: PROTEIN ENGINEERING
影响因子: --
作者:
Miyazawa, S
通讯作者: Miyazawa, S
EMBL核苷酸序列数据库:2005年的发展。
DOI: 10.1093/nar/gkj130
发表时间: 2006-01-01
影响因子: 14.9
作者:
Cochrane, Guy;Aldebert, Philippe;Althorpe, Nicola;Andersson, Mikael;Baker, Wendy;Baldwin, Alastair;Bates, Kirsty;Bhattacharyya, Sumit;Browne, Paul;van den Broek, Alexandra;Castro, Matias;Duggan, Karyn;Eberhardt, Ruth;Faruque, Nadeem;Gamble, John;Kanz, Carola;Kulikova, Tamara;Lee, Charles;Leinonen, Rasko;Lin, Quan;Lombard, Vincent;Lopez, Rodrigo;McHale, Michelle;McWilliam, Hamish;Mukherjee, Gaurab;Nardone, Francesco;Pastor, Maria Pilar Garcia;Sobhany, Siamak;Stoehr, Peter;Tzouvara, Katerina;Vaughan, Robert;Wu, Dan;Zhu, Weimin;Apweiler, Rolf
通讯作者: Apweiler, Rolf
DOI: 10.1101/gr.5890907
发表时间: 2007-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Freyhult, Eva K.;Bollback, Jonathan P.;Gardner, Paul P.
通讯作者: Gardner, Paul P.
DOI: 10.1002/pro.5560040613
发表时间: 1995-06-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
PEARSON, WR
通讯作者: PEARSON, WR