Informatics for protein identification by mass spectrometry

Informatics for protein identification by mass spectrometry
复制标题

DOI:
10.1016/j.ymeth.2004.08.014
复制
发表时间:
2005-03-01
期刊:
影响因子:
4.8
通讯作者:
Patterson, SD
Patterson, SD
中科院分区:
生物学3区
文献类型:
--
作者:
Johnson, RS;Davis, MT;Patterson, SD

文献摘要

被引文献

相似文献

高通量蛋白质分析(即,蛋白质组学)首次成为可能时,灵敏的肽质量作图技术的发展,从而允许识别和编目大多数二维凝胶电泳点的可能性。此后不久,一些研究小组率先提出了通过使用肽串联质谱搜索蛋白质序列数据库来识别蛋白质的想法。因此,从非常复杂的混合物中识别蛋白质成为可能。这些后一种技术的一个缺点是,使用修饰的肽或具有与正在搜索的序列数据库中存在的序列略有不同的序列的肽的串联质谱进行匹配并不完全简单。这是自动从头测序程序背后的动机的一部分,该程序试图推导肽序列,而不管其是否存在于序列数据库中。然后对由此产生的候选序列进行基于同源性的数据库搜索程序(例如,BLAST或FASTA)。然而,这些同源性搜索程序并没有考虑到质谱法而开发,并且有必要进行微小的修改,使得在比较查询和数据库序列时可以考虑质谱模糊性。最后,本文将讨论验证蛋白质鉴定的重要问题。所有的搜索程序都会产生一个排名靠前的答案;然而,只有轻信者才愿意接受他们的全权委托。(c)2004爱思唯尔公司All rights reserved.
High throughput protein analysis (i.e., proteomics) first became possible when sensitive peptide mass mapping techniques were developed, thereby allowing for the possibility of identifying and cataloging most 2D gel electrophoresis spots. Shortly thereafter a few groups pioneered the idea of identifying proteins by using peptide tandem mass spectra to search protein sequence databases. Hence, it became possible to identify proteins from very complex mixtures. One drawback to these latter techniques is that it is not entirely straightforward to make matches using tandem mass spectra of peptides that are modified or have sequences that differ slightly from what is present in the sequence database that is being searched. This has been part of the motivation behind automated de novo sequencing programs that attempt to derive a peptide sequence regardless of its presence in a sequence database. The sequence candidates thus generated are then subjected to homology-based database search programs (e.g., BLAST or FASTA). These homology search programs, however, were not developed with mass spectrometry in mind, and it became necessary to make minor modifications such that mass spectrometric ambiguities can be taken into account when comparing query and database sequences. Finally, this review will discuss the important issue of validating protein identifications. All of the search programs will produce a top ranked answer; however, only the credulous are willing to accept them carte blanche. (c) 2004 Elsevier Inc. All rights reserved.