Fast and accurate clustering of noncoding RNAs using ensembles of sequence alignments and secondary structures.

Fast and accurate clustering of noncoding RNAs using ensembles of sequence alignments and secondary structures.
复制标题

DOI:
10.1186/1471-2105-12-s1-s48
复制
发表时间:
2011-02-15
期刊:
影响因子:
3
通讯作者:
Sakakibara Y
Sakakibara Y
中科院分区:
生物学4区
文献类型:
--
作者:
Saito Y;Sato K;Sakakibara Y

文献摘要

被引文献

相似文献

无注释转录本的聚类是鉴定新型非编码rna (ncRNAs)家族的重要任务。基于结构对齐分数的相似性度量,已经开发了几种分层聚类方法。然而,精确结构对准的计算成本高,要求这些方法采用近似算法。这种启发式方法降低了聚类结果的质量,特别是当家庭成员之间的相似性在初级序列水平上无法检测到时。我们描述了ncrna的分层聚类的一种新的相似性度量。其思想是,通过在其动态规划框架中利用次优解的信息,可以提高近似算法的可靠性。我们以比现有方法更简化的方式近似结构对准。相反,我们的方法利用了所有可能的序列比对和所有可能的二级结构,而现有的方法只使用一个最优序列比对和一个最优二级结构。结果表明,该策略可以在计算成本和聚类质量之间达到最佳平衡。特别是当家族成员的序列识别率低于60%时,该方法仍能保持较高的性能。我们的方法能够快速准确地聚类ncrna。该软件可从http://bpla-kernel.dna.bio.keio.ac.jp/clustering/下载。
Clustering of unannotated transcripts is an important task to identify novel families of noncoding RNAs (ncRNAs). Several hierarchical clustering methods have been developed using similarity measures based on the scores of structural alignment. However, the high computational cost of exact structural alignment requires these methods to employ approximate algorithms. Such heuristics degrade the quality of clustering results, especially when the similarity among family members is not detectable at the primary sequence level. We describe a new similarity measure for the hierarchical clustering of ncRNAs. The idea is that the reliability of approximate algorithms can be improved by utilizing the information of suboptimal solutions in their dynamic programming frameworks. We approximate structural alignment in a more simplified manner than the existing methods. Instead, our method utilizes all possible sequence alignments and all possible secondary structures, whereas the existing methods only use one optimal sequence alignment and one optimal secondary structure. We demonstrate that this strategy can achieve the best balance between the computational cost and the quality of the clustering. In particular, our method can keep its high performance even when the sequence identity of family members is less than 60%. Our method enables fast and accurate clustering of ncRNAs. The software is available for download at http://bpla-kernel.dna.bio.keio.ac.jp/clustering/.