A Sensitive and Accurate protein domain cLassification Tool (SALT) for short reads

A Sensitive and Accurate protein domain cLassification Tool (SALT) for short reads
复制标题

DOI:
10.1093/bioinformatics/btt357
复制
发表时间:
2013-09-01
期刊:
影响因子:
5.8
通讯作者:
Cole, James R.
Cole, James R.
中科院分区:
生物学3区
文献类型:
--
作者:
Zhang, Yuan;Sun, Yanni;Cole, James R.

文献摘要

被引文献

相似文献

动机:蛋白质结构域分类是下一代测序数据功能注释的重要一步。对于缺乏质量或完整参考基因组的非模式生物的RNA-Seq数据,现有的蛋白质结构域分析管道被直接应用于短读取或使用从头序列组装工具产生的重叠群。结果:介绍了基于轮廓隐马尔可夫模型和图论算法的蛋白质域分类工具SALT。SALT仔细地结合了从结构域区域测序的读数的特征,并基于有监督的图构建算法将它们组装成重叠群。我们将SALT应用于两个不同读取长度的RNA-Seq数据集,并使用可用的蛋白质结构域注释和参考基因组来量化其性能。与现有策略相比,SALT显示出更高的灵敏度和准确性。在第三个实验中,我们将盐应用于非模型生物体。实验结果表明,与其他被测分类器相比,该方法能够识别出更多的转录蛋白结构域家族。
Motivation: Protein domain classification is an important step in functional annotation for next-generation sequencing data. For RNA-Seq data of non-model organisms that lack quality or complete reference genomes, existing protein domain analysis pipelines are applied to short reads directly or to contigs that are generated using de novo sequence assembly tools. However, these strategies do not provide satisfactory performance in classifying short reads into their native domain families.Results: We introduce SALT, a protein domain classification tool based on profile hidden Markov models and graph algorithms. SALT carefully incorporates the characteristics of reads that are sequenced from the domain regions and assembles them into contigs based on a supervised graph construction algorithm. We applied SALT to two RNA-Seq datasets of different read lengths and quantified its performance using the available protein domain annotations and the reference genomes. Compared with existing strategies, SALT showed better sensitivity and accuracy. In the third experiment, we applied SALT to a non-model organism. The experimental results demonstrated that it identified more transcribed protein domain families than other tested classifiers.