Searching for evolutionary distant RNA homologs within genomic sequences using partition function posterior probabilities.
Searching for evolutionary distant RNA homologs within genomic sequences using partition function posterior probabilities.
复制标题
DOI:
10.1186/1471-2105-9-61
复制
发表时间:
2008-01-28
影响因子:
3
通讯作者:
Livesay DR
中科院分区:
文献类型:
--
作者:
Roshan U;Chikkagoudar S;Livesay DR
Identification of RNA homologs within genomic stretches is difficult when pairwise sequence identity is low or unalignable flanking residues are present. In both cases structure-sequence or profile/family-sequence alignment programs become difficult to apply because of unreliable RNA structures or family alignments. As such, local sequence-sequence alignment programs are frequently used instead. We have recently demonstrated that maximal expected accuracy alignments using partition function match probabilities (implemented in Probalign) are significantly better than contemporary methods on heterogeneous length protein sequence datasets, thus suggesting an affinity for local alignment. We create a pairwise RNA-genome alignment benchmark from RFAM families with average pairwise sequence identity up to 60%. Each dataset contains a query RNA aligned to a target RNA (of the same family) embedded in a genomic sequence at least 5K nucleotides long. To simulate common conditions when exact ends of an ncRNA are unknown, each query RNA has 5' and 3' genomic flanks of size 50, 100, and 150 nucleotides. We subsequently compare the error of the Probalign program (adjusted for local alignment) to the commonly used local alignment programs HMMER, SSEARCH, and BLAST, and the popular ClustalW program with zero end-gap penalties. Parameters were optimized for each program on a small subset of the benchmark. Probalign has overall highest accuracies on the full benchmark. It leads by 10% accuracy over SSEARCH (the next best method) on 5 out of 22 families. On datasets restricted to maximum of 30% sequence identity, Probalign's overall median error is 71.2% vs. 83.4% for SSEARCH (P-value < 0.05). Furthermore, on these datasets Probalign leads SSEARCH by at least 10% on five families; SSEARCH leads Probalign by the same margin on two of the fourteen families. We also demonstrate that the Probalign mean posterior probability, compared to the normalized SSEARCH Z-score, is a better discriminator of alignment quality. All datasets and software are available online. We demonstrate, for the first time, that partition function match probabilities used for expected accuracy alignment, as done in Probalign, provide statistically significant improvement over current approaches for identifying distantly related RNA sequences in larger genomic segments.
登录
查看更多内容
DOI:
10.1016/s0097-8485(96)80004-0
发表时间:
1996-03-01
期刊:
COMPUTERS & CHEMISTRY
影响因子:
--
作者:
Gribskov, M;Robinson, NL
通讯作者:
Robinson, NL
DOI:
10.1093/protein/8.10.999
发表时间:
1995-10-01
期刊:
PROTEIN ENGINEERING
影响因子:
--
作者:
Miyazawa, S
通讯作者:
Miyazawa, S
影响因子:
14.9
作者:
Cochrane, Guy;Aldebert, Philippe;Althorpe, Nicola;Andersson, Mikael;Baker, Wendy;Baldwin, Alastair;Bates, Kirsty;Bhattacharyya, Sumit;Browne, Paul;van den Broek, Alexandra;Castro, Matias;Duggan, Karyn;Eberhardt, Ruth;Faruque, Nadeem;Gamble, John;Kanz, Carola;Kulikova, Tamara;Lee, Charles;Leinonen, Rasko;Lin, Quan;Lombard, Vincent;Lopez, Rodrigo;McHale, Michelle;McWilliam, Hamish;Mukherjee, Gaurab;Nardone, Francesco;Pastor, Maria Pilar Garcia;Sobhany, Siamak;Stoehr, Peter;Tzouvara, Katerina;Vaughan, Robert;Wu, Dan;Zhu, Weimin;Apweiler, Rolf
通讯作者:
Apweiler, Rolf
影响因子:
7
作者:
Freyhult, Eva K.;Bollback, Jonathan P.;Gardner, Paul P.
通讯作者:
Gardner, Paul P.
影响因子:
8
作者:
PEARSON, WR
通讯作者:
PEARSON, WR