Annotation of metagenome short reads using proxygenes

Annotation of metagenome short reads using proxygenes
复制标题

DOI:
10.1093/bioinformatics/btn276
复制
发表时间:
2008-08-15
期刊:
影响因子:
5.8
通讯作者:
Markowitz, Victor M.
Markowitz, Victor M.
中科院分区:
生物学3区
文献类型:
--
作者:
Dalevi, Daniel;Ivanova, Natalia N.;Markowitz, Victor M.

文献摘要

被引文献

相似文献

动机:使用454焦磷酸测序平台生成的典型宏基因组数据集由从微生物群落的集体基因组采样的短读段组成。这样的数据集中的序列量通常不足以进行组装,并且传统的基因预测不能应用于未组装的短读段。因此,对这些数据集的分析通常涉及各种蛋白质家族的相对丰度的比较。后者需要分配的个人阅读的蛋白质家族,这是阻碍了短的阅读只包含一个片段,通常很小,一个protein.Results的事实:我们已经考虑了分配焦磷酸测序读到蛋白质家族直接使用RPS-BLAST对COG和Pfam数据库和间接通过proxygenes确定使用BLASTx搜索对蛋白质序列数据库。使用模拟宏基因组数据集作为基准,我们表明,proxygene方法比直接分配更准确。我们介绍了一种聚类方法,它显着降低了宏基因组数据集的大小,同时保持其功能和分类内容的忠实代表。
Motivation: A typical metagenome dataset generated using a 454 pyrosequencing platform consists of short reads sampled from the collective genome of a microbial community. The amount of sequence in such datasets is usually insufficient for assembly, and traditional gene prediction cannot be applied to unassembled short reads. As a result, analysis of such datasets usually involves comparisons in terms of relative abundances of various protein families. The latter requires assignment of individual reads to protein families, which is hindered by the fact that short reads contain only a fragment, usually small, of a protein.Results: We have considered the assignment of pyrosequencing reads to protein families directly using RPS-BLAST against COG and Pfam databases and indirectly via proxygenes that are identified using BLASTx searches against protein sequence databases. Using simulated metagenome datasets as benchmarks, we show that the proxygene method is more accurate than the direct assignment. We introduce a clustering method which significantly reduces the size of a metagenome dataset while maintaining a faithful representation of its functional and taxonomic content.