Simultaneous gene finding in multiple genomes

Simultaneous gene finding in multiple genomes
复制标题

DOI:
10.1093/bioinformatics/btw494
复制
发表时间:
2016-11-15
期刊:
影响因子:
5.8
通讯作者:
Stanke, Mario
Stanke, Mario
中科院分区:
生物学3区
文献类型:
--
作者:
Koenig, Stefanie;Romoth, Lars W.;Stanke, Mario

文献摘要

被引文献

相似文献

动机:随着生命之树被越来越密集的测序基因组所填充,新的挑战是对整个基因组分支的准确和一致的注释。我们通过一种新的比较基因发现方法来解决这个问题,该方法通过对密切相关物种的多个基因组进行比对,同时预测所有输入基因组中编码蛋白质的基因的位置和结构,从而利用负选择和序列保护。该模型倾向于不同基因组中彼此一致的潜在基因结构,或者--如果不一致--在给定物种树的情况下,外显子的得失是可信的。我们将多物种基因发现问题描述为图上的二进制标记问题。所得到的优化问题是NP困难的,但可以使用基于次梯度的对偶分解方法有效地逼近。结果:所提出的方法在12种脊椎动物和12种果蝇的全基因组比对中得到了测试。对人类、小鼠和果蝇的准确度进行了评估,并与竞争方法进行了比较。结果表明,我们的方法非常适合于注释一个分支中密切相关物种的(大量)基因组,特别是当许多基因组有RNA-Seq数据时。当基因组接近中等距离时,通过基因组比对将现有注释从一个基因组转移到另一个基因组比基于蛋白质剪接比对的以前的方法更准确。
Motivation: As the tree of life is populated with sequenced genomes ever more densely, the new challenge is the accurate and consistent annotation of entire clades of genomes. We address this problem with a new approach to comparative gene finding that takes a multiple genome alignment of closely related species and simultaneously predicts the location and structure of protein-coding genes in all input genomes, thereby exploiting negative selection and sequence conservation. The model prefers potential gene structures in the different genomes that are in agreement with each other, or-if not-where the exon gains and losses are plausible given the species tree. We formulate the multi-species gene finding problem as a binary labeling problem on a graph. The resulting optimization problem is NP hard, but can be efficiently approximated using a subgradient-based dual decomposition approach.Results: The proposed method was tested on whole-genome alignments of 12 vertebrate and 12 Drosophila species. The accuracy was evaluated for human, mouse and Drosophila melanogaster and compared to competing methods. Results suggest that our method is well-suited for annotation of (a large number of) genomes of closely related species within a clade, in particular, when RNA-Seq data are available for many of the genomes. The transfer of existing annotations from one genome to another via the genome alignment is more accurate than previous approaches that are based on protein-spliced alignments, when the genomes are at close to medium distances.