Mass spectrometry-based prokaryote gene annotation

Mass spectrometry-based prokaryote gene annotation
复制标题

DOI:
10.1002/pmic.200700080
复制
发表时间:
2007-11-01
期刊:
影响因子:
3.4
通讯作者:
Taniguchi, Hisaaki
Taniguchi, Hisaaki
中科院分区:
生物学3区
文献类型:
--
作者:
Ishino, Yoko;Okada, Hitomi;Taniguchi, Hisaaki

文献摘要

被引文献

相似文献

MS结合数据库检索已成为鉴定细胞或组织样品中存在的蛋白质的首选方法。这项技术使我们能够对已经测序的物种进行大规模的蛋白质组分析。针对由注释基因组成的蛋白质数据库进行质谱数据检索已被广泛开展。然而,这种技术存在一些问题;蛋白质数据库中错误的注释会导致蛋白质鉴定的准确性下降,只有已经被注释过的蛋白质才能被鉴定出来。我们提出了一个新的框架,可以通过整合MS/MS蛋白质组学数据映射和基于知识的翻译起始位点系统来检测正确的orf。该技术可以提供预测的编码序列的校正,以及鉴定新基因的可能性。我们开发了一个计算系统;首先利用MS/MS数据对所有可能的翻译框进行概率多肽匹配,然后在检测到的多肽周围搜索有区别的DNA模式,最后利用知识库中存储的经验知识对事实进行整合,得到正确的orf。我们以光合细菌Synechocystis sp. PCC6803为原核生物样本,发现了14个n端注释错误和几个新的候选基因。
MS combined with database searching has become the preferred method for identifying proteins present in cell or tissue samples. The technique enables us to execute large-scale proteome analyses of species whose genomes have already been sequenced. Searching mass spectrometric data against protein databases composed of annotated genes has been widely conducted. However, there are some issues with this technique; wrong annotations in protein databases cause deterioration in the accuracy of protein identification, and only proteins that have already been annotated can be identified. We propose a new framework that can detect correct ORFs by integrating an MS/MS proteomic data mapping and a knowledge-based system regarding the translation initiation sites. This technique can provide correction of predicted coding sequences, together with the possibility of identifying novel genes. We have developed a computational system; it should first conduct the probabilistic peptide-matching against all possible translational frames using MS/MS data, then search for discriminative DNA patterns around the detected peptides, and lastly integrate the facts using empirical knowledge stored in knowledge bases to obtain correct ORFs. We used photosynthetic bacteria Synechocystis sp. PCC6803 as a sample prokaryote, resulting in the finding of 14 N-terminus annotation errors and several new candidate genes.