Integrated approach for manual evaluation of peptides identified by searching protein sequence databases with tandem mass spectra

Integrated approach for manual evaluation of peptides identified by searching protein sequence databases with tandem mass spectra
复制标题

DOI:
10.1021/pr049754t
复制
发表时间:
2005-05-01
影响因子:
4.4
通讯作者:
Zhao, YM
Zhao, YM
中科院分区:
生物学2区
文献类型:
--
作者:
Chen, Y;Kwon, SW;Zhao, YM

文献摘要

被引文献

相似文献

定量蛋白质组学依赖于准确的蛋白质鉴定,这通常是通过自动搜索序列数据库与串联质谱的肽。当这些光谱包含有限的信息时,自动搜索可能导致不正确的肽鉴定。因此,有必要通过仔细手动检查质谱来验证鉴别。这项任务不仅耗时,而且验证的可靠性也随分析人员的经验而变化。在这里,我们报告了一个系统的方法来评估自动搜索算法的肽鉴定。该方法基于候选肽序列应充分解释观察到的碎片离子的原则。此外,相邻碎片的质量误差应该相似。为了评价我们的方法,我们研究了从E. coli和HeLa细胞。用自动搜索引擎Mascot鉴定候选肽,并进行手动验证方法。该方法发现了正确的肽鉴定,这些肽被给予低Mascot分数(例如,20-25)和给予高Mascot分数的不正确的肽鉴定(例如,40-50)。该方法全面地检测了旨在产生不正确识别的搜索的错误结果。合成候选肽的串联质谱与从复杂肽混合物获得的光谱的比较证实了评价方法的准确度。因此,本文所述的评估方法可以帮助提高蛋白质鉴定的准确性,增加鉴定的肽的数量,并为开发更准确的下一代蛋白质鉴定算法迈出了一步。
Quantitative proteomics relies on accurate protein identification, which often is carried out by automated searching of a sequence database with tandem mass spectra of peptides. When these spectra contain limited information, automated searches may lead to incorrect peptide identifications. It is therefore necessary to validate the identifications by careful manual inspection of the mass spectra. Not only is this task time-consuming, but the reliability of the validation varies with the experience of the analyst. Here, we report a systematic approach to evaluating peptide identifications made by automated search algorithms. The method is based on the principle that the candidate peptide sequence should adequately explain the observed fragment ions. Also, the mass errors of neighboring fragments should be similar. To evaluate our method, we studied tandem mass spectra obtained from tryptic digests of E. coli and HeLa cells. Candidate peptides were identified with the automated search engine Mascot and subjected to the manual validation method. The method found correct peptide identifications that were given low Mascot scores (e.g., 20-25) and incorrect peptide identifications that were given high Mascot scores (e.g., 40-50). The method comprehensively detected false results from searches designed to produce incorrect identifications. Comparison of the tandem mass spectra of synthetic candidate peptides to the spectra obtained from the complex peptide mixtures confirmed the accuracy of the evaluation method. Thus, the evaluation approach described here could help boost the accuracy of protein identification, increase number of peptides identified, and provide a step toward developing a more accurate next-generation algorithm for protein identification.