FragGeneScan: predicting genes in short and error-prone reads.

FragGeneScan: predicting genes in short and error-prone reads.
复制标题

DOI:
10.1093/nar/gkq747
复制
发表时间:
2010-11
影响因子:
14.9
通讯作者:
Ye Y
Ye Y
中科院分区:
生物学2区
文献类型:
--
作者:
Rho M;Tang H;Ye Y

文献摘要

参考文献

被引文献

相似文献

下一代测序技术的进步促进了宏基因组学研究,该研究试图直接确定环境样品中的整个遗传物质集合(即宏基因组)。由于宏基因组的组装通常不可用,因此直接从短读段中鉴定基因已经成为注释宏基因组中重要但具有挑战性的问题。为全基因组开发的基因预测器(例如Glimmer)和最近为宏基因组序列开发的基因预测器(例如MetaGene)显示出随着测序错误率增加或随着读数变短而性能显著降低。我们已经开发了一种新的基因预测方法FragGeneScan,它结合了测序错误模型和密码子使用的隐马尔可夫模型,以提高预测蛋白质编码区在短读段。对于完整基因组,FragGeneScan的性能与Glimmer和MetaGene相当。但对于短读段,FragGeneScan始终优于MetaGene(对于400个碱基的读段,准确性提高了62%,测序错误率为1%,对于100个碱基的无错误短读段,准确性提高了18%)。当应用于宏基因组时,FragGeneScan恢复了比MetaGene预测的更多的基因(>90%的基因通过同源性搜索被鉴定),并且许多新基因在当前蛋白质序列数据库中没有同源物。
The advances of next-generation sequencing technology have facilitated metagenomics research that attempts to determine directly the whole collection of genetic material within an environmental sample (i.e. the metagenome). Identification of genes directly from short reads has become an important yet challenging problem in annotating metagenomes, since the assembly of metagenomes is often not available. Gene predictors developed for whole genomes (e.g. Glimmer) and recently developed for metagenomic sequences (e.g. MetaGene) show a significant decrease in performance as the sequencing error rates increase, or as reads get shorter. We have developed a novel gene prediction method FragGeneScan, which combines sequencing error models and codon usages in a hidden Markov model to improve the prediction of protein-coding region in short reads. The performance of FragGeneScan was comparable to Glimmer and MetaGene for complete genomes. But for short reads, FragGeneScan consistently outperformed MetaGene (accuracy improved ∼62% for reads of 400 bases with 1% sequencing errors, and ∼18% for short reads of 100 bases that are error free). When applied to metagenomes, FragGeneScan recovered substantially more genes than MetaGene predicted (>90% of the genes identified by homology search), and many novel genes with no homologs in current protein sequence database.
DOI: 10.1093/dnares/dsn027
发表时间: 2008-12
期刊: DNA RESEARCH
影响因子: 4.1
作者:
Noguchi, Hideki;Taniguchi, Takeaki;Itoh, Takehiko
通讯作者: Itoh, Takehiko
DOI: 10.1186/gb-2009-10-8-r83
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Kircher M;Stenzel U;Kelso J
通讯作者: Kelso J
DOI: 10.1126/science.1124234
发表时间: 2006-06-02
期刊: SCIENCE
影响因子: 56.9
作者:
Gill, Steven R.;Pop, Mihai;Nelson, Karen E.
通讯作者: Nelson, Karen E.
DOI: 10.1186/gb-2008-9-2-r41
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Hu G;Liu I;Sham A;Stajich JE;Dietrich FS;Kronstad JW
通讯作者: Kronstad JW
DOI: 10.1093/nar/gkl723
发表时间: 2006-11-01
影响因子: 14.9
作者:
Noguchi, Hideki;Park, Jungho;Takagi, Toshihisa
通讯作者: Takagi, Toshihisa