Sequence-based heuristics for faster annotation of non-coding RNA families

Sequence-based heuristics for faster annotation of non-coding RNA families
复制标题

DOI:
10.1093/bioinformatics/bti743
复制
发表时间:
2006-01-01
期刊:
影响因子:
5.8
通讯作者:
Ruzzo, WL
Ruzzo, WL
中科院分区:
生物学3区
文献类型:
--
作者:
Weinberg, Z;Ruzzo, WL

文献摘要

被引文献

相似文献

非编码RNA(ncRNA)是不编码蛋白质的功能性RNA分子。协方差模型(CM)是一个有用的统计工具,发现新成员的ncRNA基因家族在一个大型的基因组数据库中,使用序列和RNA的二级结构信息,重要的是。不幸的是,CM搜索非常慢。在此之前,我们创建了严格的过滤器,可证明不会牺牲CM的准确性,同时使几乎所有ncRNA家族的搜索速度显著加快。然而,这些严格的过滤器使搜索速度慢于启发式cannes.Results:在本文中,我们介绍了配置文件HMM为基础的启发式过滤器。我们表明,他们的准确性通常是上级的基础上BLAST的分析。此外,我们比较了我们的分子生物学与tRNAscan-SE中使用的分子生物学,tRNAscan-SE的分子生物学结合了大量特定于tRNA的工作,而我们的分子生物学对任何ncRNA都是通用的。性能大致相当,因此我们希望我们的免疫学提供一种高质量的解决方案,与家族特异性解决方案不同,可以扩展到数百个ncRNA家族。
Motivation: Non-coding RNAs (ncRNAs) are functional RNA molecules that do not code for proteins. Covariance Models (CMs) are a useful statistical tool to find new members of an ncRNA gene family in a large genome database, using both sequence and, importantly, RNA secondary structure information. Unfortunately, CM searches are extremely slow. Previously, we created rigorous filters, which provably sacrifice none of a CM's accuracy, while making searches significantly faster for virtually all ncRNA families. However, these rigorous filters make searches slower than heuristics could be.Results: In this paper we introduce profile HMM-based heuristic filters. We show that their accuracy is usually superior to heuristics based on BLAST. Moreover, we compared our heuristics with those used in tRNAscan-SE, whose heuristics incorporate a significant amount of work specific to tRNAs, where our heuristics are generic to any ncRNA. Performance was roughly comparable, so we expect that our heuristics provide a high-quality solution that-unlike family-specific solutions-can scale to hundreds of ncRNA families.