Performance and scalability of discriminative metrics for comparative gene identification in 12 Drosophila genomes.

Performance and scalability of discriminative metrics for comparative gene identification in 12 Drosophila genomes.
复制标题

在12个果蝇基因组中进行比较基因鉴定的判别指标的性能和可伸缩性。

DOI:
10.1371/journal.pcbi.1000067
复制
发表时间:
2008-04-18
影响因子:
4.3
通讯作者:
Kellis, Manolis
Kellis, Manolis
中科院分区:
生物学2区
文献类型:
--
作者:
Lin, Michael F.;Deoras, Ameya N.;Rasmussen, Matthew D.;Kellis, Manolis

文献摘要

参考文献

被引文献

相似文献

多个相关物种的比较基因组学是发现功能性基因组元件的一种强大方法,并且其效力应随着所比较的物种数量增加而提高。在此,我们利用12个果蝇基因组来研究比较基因组学指标区分蛋白质编码区和非编码区的能力。首先,我们研究了不同比较指标的相对效力及其与单物种指标的关系。我们发现,即使是相对简单的多物种指标也显著优于先进的单物种指标,尤其是对于动物基因组中常见的较短外显子(≤240个核苷酸)。此外,这两种指标在很大程度上捕捉到了蛋白质编码基因的独立特征,具有不同的敏感性/特异性权衡,因此它们的组合会产生更强的区分能力。此外,我们研究了发现能力如何随着所比较的基因组数量和系统发育距离而变化。我们发现,在较广的距离范围内的物种对于成对比较基因鉴定是同样有效的信息源,但在相似的进化分歧下,多物种比较的效果更佳。特别是,虽然成对发现能力在较大距离时达到稳定且从未优于最先进的单物种指标,但多物种比较即使在最遥远的物种中也能持续受益,且没有明显的饱和现象。最后,我们发现,通常被认为是快速进化的功能类别中的基因仍然可以使用比较方法以很高的比率被恢复。我们的研究结果对任何物种(包括人类)的比较基因组学分析都具有启示意义。 比较相关物种的基因组是发现诸如蛋白质编码基因等功能元件的一种强大方法。从理论上讲,使用更多的物种应该会带来更强的发现能力。然而,关于比较哪些物种是最优选择以及如何最好地利用多物种比对,仍然存在许多问题。甚至有可能基因组测序、组装和比对的实际限制会有效地抵消使用更多物种的益处。在此,我们使用12个完整的果蝇基因组来研究用于鉴定蛋白质编码基因的各种指标,包括仅分析目标基因组的方法以及检查基因组比对中进化特征的比较方法。我们发现,在令人惊讶的较广系统发育距离范围内的物种在比较分析中是有效的,并且发现能力随着每增加一个物种而持续提高,且没有明显的饱和现象。我们还研究了比较方法是否会系统性地遗漏被认为是快速进化的基因,并研究了基因组比对策略如何影响性能。我们的结果可以帮助指导未来比较研究的物种选择,并为各种基因鉴定任务提供方法学指导,包括未来从头基因预测器的设计以及对特殊基因结构的搜索。
Comparative genomics of multiple related species is a powerful methodology for the discovery of functional genomic elements, and its power should increase with the number of species compared. Here, we use 12 Drosophila genomes to study the power of comparative genomics metrics to distinguish between protein-coding and non-coding regions. First, we study the relative power of different comparative metrics and their relationship to single-species metrics. We find that even relatively simple multi-species metrics robustly outperform advanced single-species metrics, especially for shorter exons (≤240 nt), which are common in animal genomes. Moreover, the two capture largely independent features of protein-coding genes, with different sensitivity/specificity trade-offs, such that their combinations lead to even greater discriminatory power. In addition, we study how discovery power scales with the number and phylogenetic distance of the genomes compared. We find that species at a broad range of distances are comparably effective informants for pairwise comparative gene identification, but that these are surpassed by multi-species comparisons at similar evolutionary divergence. In particular, while pairwise discovery power plateaued at larger distances and never outperformed the most advanced single-species metrics, multi-species comparisons continued to benefit even from the most distant species with no apparent saturation. Last, we find that genes in functional categories typically considered fast-evolving can nonetheless be recovered at very high rates using comparative methods. Our results have implications for comparative genomics analyses in any species, including the human. Comparing the genomes of related species is a powerful approach to the discovery of functional elements such as protein-coding genes. Theoretically, using more species should lead to more discovery power. Many questions remain, however, surrounding the optimal choice of species to compare and how to best use multi-species alignments. It is even possible that practical limitations in the sequencing, assembly, and alignment of genomes could effectively negate the benefit of using more species. Here, we used 12 complete fly genomes to study a variety of metrics used to identify protein-coding genes, including methods that analyze only the genome of interest and comparative methods that examine evolutionary signatures in genome alignments. We found that species over a surprisingly broad range of phylogenetic distances were effective in comparative analyses, and that discovery power continued to scale with each additional species without apparent saturation. We also examined whether comparative methods systematically miss genes considered fast-evolving, and studied how performance is influenced by genome alignment strategies. Our results can help guide species selection for future comparative studies and provide methodological guidance for a variety of gene identification tasks, including the design of future de novo gene predictors and the search for unusual gene structures.
DOI: 10.1038/nature06341
发表时间: 2007-11-08
期刊: NATURE
影响因子: 64.8
作者:
Clark, Andrew G.;Eisen, Michael B.;MacCallum, Iain
通讯作者: MacCallum, Iain
DOI: 10.1101/gr.3866105
发表时间: 2005-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Brent, MR
通讯作者: Brent, MR
DOI: 10.1109/79.939833
发表时间: 2001-07-01
影响因子: 14.9
作者:
Anastassiou, D
通讯作者: Anastassiou, D
DOI: 10.1371/journal.pbio.0030010
发表时间: 2005-01-01
期刊: PLOS BIOLOGY
影响因子: 9.8
作者:
Eddy, SR
通讯作者: Eddy, SR
DOI: 10.1101/gr.424203
发表时间: 2003-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Alexandersson, M;Cawley, S;Pachter, L
通讯作者: Pachter, L