Ortho-proteogenomics: Multiple proteomes investigation through orthology and a new MS-based protocol

Ortho-proteogenomics: Multiple proteomes investigation through orthology and a new MS-based protocol
复制标题

DOI:
10.1101/gr.081901.108
复制
发表时间:
2009-01-01
期刊:
影响因子:
7
通讯作者:
Lecompte, Odile
Lecompte, Odile
中科院分区:
生物学1区
文献类型:
--
作者:
Gallien, Sebastien;Perrodou, Emmanuel;Lecompte, Odile

文献摘要

被引文献

相似文献

测序技术的进步为生物学提供了越来越多的基因组序列。在大多数情况下,基因谱系是在电子计算机中预测的,并在概念上翻译成蛋白质。正如最近强调的那样,预测的基因经常出现错误,特别是在起始密码子中,这对后续的生物学研究产生了严重影响。本文提出了一种新的“正蛋白质组学”方法,用于同时对多个基因组进行注释精化。它结合了比较基因组学和原始的蛋白质组学方案,允许在单个实验中同时表征N末端和内部多肽。以污垢分枝杆菌为参照,将该策略应用于分枝杆菌属,共鉴定出946个差异蛋白质,其中包括443个特征N末端。这些实验数据纠正了19%的特征起始密码子,鉴定了注释过程中遗漏的29种蛋白质,并通过比较基因组学对16个其他分枝杆菌蛋白质组的4328个序列进行了整理。
The progress in sequencing technologies irrigates biology with an ever-increasing number of genome sequences. In most cases, the gene repertoire is predicted in silico and conceptually translated into proteins. As recently highlighted, the predicted genes exhibit frequent errors, particularly in start codons, with a serious impact on subsequent biological studies. A new "ortho-proteogenomic" approach is presented here for the annotation refinement of multiple genomes at once. It combines comparative genomics with an original proteomic protocol that allows the characterization of both N-terminal and internal peptides in a single experiment. This strategy was applied to the Mycobacterium genus with Mycobacterium smegmatis as the reference, and identified 946 distinct proteins, including 443 characterized N termini. These experimental data allowed the correction of 19% of the characterized start codons, the identification of 29 proteins missed during the annotation process, and the curation, thanks to comparative genomics, of 4328 sequences of 16 other Mycobacterium proteomes.