Using intron position conservation for homology-based gene prediction.

Using intron position conservation for homology-based gene prediction.
复制标题

DOI:
10.1093/nar/gkw092
复制
发表时间:
2016-05-19
影响因子:
14.9
通讯作者:
Hartung F
Hartung F
中科院分区:
生物学2区
文献类型:
--
作者:
Keilwagen J;Wenk M;Erickson JL;Schattat MH;Grau J;Hartung F

文献摘要

被引文献

相似文献

蛋白质编码基因的注释在生物信息学和生物学中非常重要,对许多下游分析具有决定性影响。基于同源性的基因预测程序允许将关于蛋白质编码基因的知识从注释的生物体转移到感兴趣的生物体。在这里,我们提出了一个同源性为基础的基因预测程序称为GeMoMa。GeMoMa利用基因内内含子位置的保守性来预测其他生物中的相关基因。我们评估了GeMoMa的性能,并使用扩展的最佳互惠命中方法将其与植物和动物基因组上最先进的竞争对手进行比较。我们发现,GeMoMa往往比它的竞争对手更精确的预测产生大量增加的正确成绩单。随后,我们使用桑格测序示范性地验证GeMoMa预测。最后,我们使用RNA-seq数据来比较基于同源性的基因预测程序的预测,并再次发现GeMoMa表现良好。因此,我们得出结论,利用内含子位置保守性可以提高基于同源性的基因预测,并且我们将GeMoMa作为命令行工具和Galaxy集成免费提供。
Annotation of protein-coding genes is very important in bioinformatics and biology and has a decisive influence on many downstream analyses. Homology-based gene prediction programs allow for transferring knowledge about protein-coding genes from an annotated organism to an organism of interest. Here, we present a homology-based gene prediction program called GeMoMa. GeMoMa utilizes the conservation of intron positions within genes to predict related genes in other organisms. We assess the performance of GeMoMa and compare it with state-of-the-art competitors on plant and animal genomes using an extended best reciprocal hit approach. We find that GeMoMa often makes more precise predictions than its competitors yielding a substantially increased number of correct transcripts. Subsequently, we exemplarily validate GeMoMa predictions using Sanger sequencing. Finally, we use RNA-seq data to compare the predictions of homology-based gene prediction programs, and find again that GeMoMa performs well. Hence, we conclude that exploiting intron position conservation improves homology-based gene prediction, and we make GeMoMa freely available as command-line tool and Galaxy integration.