Robust prediction of consensus secondary structures using averaged base pairing probability matrices

Robust prediction of consensus secondary structures using averaged base pairing probability matrices
复制标题

DOI:
10.1093/bioinformatics/btl636
复制
发表时间:
2007-02-15
期刊:
影响因子:
5.8
通讯作者:
Asai, Kiyoshi
Asai, Kiyoshi
中科院分区:
生物学3区
文献类型:
--
作者:
Kiryu, Hisanori;Kin, Taishin;Asai, Kiyoshi

文献摘要

被引文献

相似文献

动机:最近的转录组学研究表明,在高等真核细胞中存在相当数量的非蛋白编码RNA转录物。为了研究这些转录本的功能作用,在基因组尺度上从多个序列中发现保守的二级结构是非常有趣的。由于经常使用为了计算效率而忽略RNA二级结构特殊保守模式的比对程序创建多个比对,因此比对失败可能会导致忽略保守茎结构的潜在风险。结果:我们研究了二级结构预测精度与牙体矫正质量的关系。我们比较了三种算法,最大限度地提高了二级结构的预期精度,以及其他常用的算法。我们发现我们的一种算法,称为McCaskill-MEA,比其他算法对对齐失败更健壮。McCaskill-MEA方法首先计算比对中所有序列的碱基配对概率矩阵,然后对这些矩阵进行平均,得到比对的碱基配对概率矩阵。从该矩阵预测共识二级结构,使预测的预期精度最大化。结果表明,McCaskill-MEA方法在比对质量较低和比对由多个序列组成的情况下比其他方法性能更好。我们的模型有一个参数来控制预测的敏感性和特异性。我们讨论了该参数在多步筛选过程中的使用,以搜索保守的二级结构,并为预测的碱基对分配置信度值。可用性:实现McCaskill-MEA算法的c++源代码和本文使用的测试数据集可在http://www.ncrna.org/papers/McCaskillMEA/Contact: kiryu-h@aist.go.jpSupplementary上获得信息:补充数据可在Bioinformatics online上获得。
Motivation: Recent transcriptomic studies have revealed the existence of a considerable number of non-protein-coding RNA transcripts in higher eukaryotic cells. To investigate the functional roles of these transcripts, it is of great interest to find conserved secondary structures from multiple alignments on a genomic scale. Since multiple alignments are often created using alignment programs that neglect the special conservation patterns of RNA secondary structures for computational efficiency, alignment failures can cause potential risks of overlooking conserved stem structures.Results: We investigated the dependence of the accuracy of secondary structure prediction on the quality of alignments. We compared three algorithms that maximize the expected accuracy of secondary structures as well as other frequently used algorithms. We found that one of our algorithms, called McCaskill-MEA, was more robust against alignment failures than others. The McCaskill-MEA method first computes the base pairing probability matrices for all the sequences in the alignment and then obtains the base pairing probability matrix of the alignment by averaging over these matrices. The consensus secondary structure is predicted from this matrix such that the expected accuracy of the prediction is maximized. We show that the McCaskill-MEA method performs better than other methods, particularly when the alignment quality is low and when the alignment consists of many sequences. Our model has a parameter that controls the sensitivity and specificity of predictions. We discussed the uses of that parameter for multi-step screening procedures to search for conserved secondary structures and for assigning confidence values to the predicted base pairs.Availability: The C++ source code that implements the McCaskill-MEA algorithm and the test dataset used in this paper are available at http://www.ncrna.org/papers/McCaskillMEA/Contact: kiryu-h@aist.go.jpSupplementary information: Supplementary data are available at Bioinformatics online.