Proteogenomics to discover the full coding content of genomes: a computational perspective.

Proteogenomics to discover the full coding content of genomes: a computational perspective.
复制标题

DOI:
10.1016/j.jprot.2010.06.007
复制
发表时间:
2010-10-10
影响因子:
3.3
通讯作者:
Bafna V
Bafna V
中科院分区:
生物学2区
文献类型:
--
作者:
Castellana N;Bafna V

文献摘要

参考文献

被引文献

相似文献

蛋白质组学是基因组学和蛋白质组学交叉的一个研究领域。它是一个松散的技术集合,允许对基因组数据库进行串联质谱搜索,以识别和表征蛋白质编码基因。蛋白基因组多肽为基因注释提供了宝贵的信息,这是很难或不可能确定使用标准的注释方法。例子包括翻译的确认、阅读框的确定、基因和外显子边界的鉴定、翻译后加工的证据、包括选择性剪接的剪接形式的鉴定,以及完全新基因的预测。然而,蛋白质基因组学要实现其承诺,必须克服许多技术障碍,包括肽鉴定的速度和准确性,专业数据库的构建和搜索,采样偏差的校正等。本文回顾了该领域的最新技术,重点介绍了当前的成功,以及计算在克服这些挑战中的作用。我们描述了如何技术和算法的进步已经使大规模的蛋白基因组学研究在许多模式生物,包括拟南芥,酵母,苍蝇和人类。我们还提供了一个前瞻性的领域前进,描述了早期的努力,在解决复杂的基因结构的问题,对相关物种的基因组搜索,免疫球蛋白基因重建。
Proteogenomics has emerged as a field at the junction of genomics and proteomics. It is a loose collection of technologies that allow the search of tandem mass spectra against genomic databases to identify and characterize protein-coding genes. Proteogenomic peptides provide invaluable information for gene annotation, which is difficult or impossible to ascertain using standard annotation methods. Examples include confirmation of translation, reading-frame determination, identification of gene and exon boundaries, evidence for post-translational processing, identification of splice-forms including alternative splicing, and also, prediction of completely novel genes. For proteogenomics to deliver on its promise, however, it must overcome a number of technological hurdles, including speed and accuracy of peptide identification, construction and search of specialized databases, correction of sampling bias, and others. This article reviews the state of the art of the field, focusing on the current successes, and the role of computation in overcoming these challenges. We describe how technological and algorithmic advances have already enabled large-scale proteogenomic studies in many model organisms, including arabidopsis, yeast, fly, and human. We also provide a preview of the field going forward, describing early efforts in tackling the problems of complex gene structures, searching against genomes of related species, and immunoglobulin gene reconstruction.
DOI: 10.1002/0471250953.bi0403s18
发表时间: 2007-06-01
影响因子: --
作者:
Blanco, Enrique;Parra, Genis;Guigo, Roderic
通讯作者: Guigo, Roderic
DOI: 10.1101/gr.1858004
发表时间: 2004-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Curwen, V;Eyras, E;Clamp, M
通讯作者: Clamp, M
DOI: 10.1093/bioinformatics/bth092
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Craig, R;Beavis, RC
通讯作者: Beavis, RC
DOI: 10.1038/ng.128
发表时间: 2008-06
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Campbell, Peter J.;Stephens, Philip J.;Pleasance, Erin D.;O'Meara, Sarah;Li, Heng;Santarius, Thomas;Stebbings, Lucy A.;Leroy, Catherine;Edkins, Sarah;Hardy, Claire;Teague, Jon W.;Menzies, Andrew;Goodhead, Ian;Turner, Daniel J.;Clee, Christopher M.;Quail, Michael A.;Cox, Antony;Brown, Clive;Durbin, Richard;Hurles, Matthew E.;Edwards, Paul A. W.;Bignell, Graham R.;Stratton, Michael R.;Futreal, P. Andrew
通讯作者: Futreal, P. Andrew
DOI: 10.1073/pnas.0811066106
发表时间: 2008-12-30
影响因子: 11.1
作者:
Castellana, Natalie E.;Payne, Samuel H.;Briggs, Steven P.
通讯作者: Briggs, Steven P.