Optimization of Search Engines and Postprocessing Approaches to Maximize Peptide and Protein Identification for High-Resolution Mass Data.

Optimization of Search Engines and Postprocessing Approaches to Maximize Peptide and Protein Identification for High-Resolution Mass Data.
复制标题

DOI:
10.1021/acs.jproteome.5b00536
复制
发表时间:
2015-11-06
影响因子:
4.4
通讯作者:
Qu J
Qu J
中科院分区:
生物学2区
文献类型:
--
作者:
Tu C;Sheng Q;Li J;Ma D;Shen X;Wang X;Shyr Y;Yi Z;Qu J

文献摘要

被引文献

相似文献

分析高分辨率质谱产生的蛋白质组数据的两个关键步骤是数据库搜索和后处理。虽然这两个步骤是相互关联的,但对它们的组合效应和这些程序的优化尚未进行充分的研究。在这里,我们研究了三个流行的搜索引擎(SEQUEST,Mascot和MS阿曼达)的性能与五种过滤方法,包括各自的分数为基础的过滤,基于组的方法,本地错误发现率(LFDR),PeptideProphet和Percolator。来自各种蛋白质组的总共八个数据集(例如,e.大肠杆菌、酵母菌和人)进行分析。发现对于所有数据集,涉及Percolator的组合在相同FDR水平下比其他12种组合实现了显著更多的肽和蛋白质鉴定。其中,SEQUEST-Percolator和MS Amanda-Percolator的组合分别为低精度MS2(离子阱或IT)和高精度MS2(Orbitrap或TOF)的数据集提供了比其他方法稍好的性能。对于不使用Percolator的方法,SEQUEST-group对碰撞诱导解离(CID)和IT分析产生的MS2数据集表现最好; Mascot-LFDR为高能碰撞解离(HCD)产生的数据集提供了更多识别,并在Orbitrap(HCD-OT)和Orbitrap Fusion(HCD-IT)中进行了分析; MS Amanda-Group在Q-TOF数据集和Orbitrap Velos HCD-OT数据集方面表现出色。因此,如果未使用Percolator,则应将特定组合应用于每种类型的数据集。此外,当使用Percolator相关组合分析技术重复时,观察到更高百分比的多肽蛋白和更低的蛋白质光谱计数变化;因此,Percolator增强了鉴定和定量的可靠性。使用嵌入Proteome Discoverer、Scaffold中的特定程序和内部算法(Build Summary)进行分析。这些结果为蛋白质组学结果的最佳解释和在不同情况下开发适合目的的协议提供了有价值的指导方针。
The two key steps for analyzing proteomic data generated by high-resolution MS are database searching and postprocessing. While the two steps are interrelated, studies on their combinatory effects and the optimization of these procedures have not been adequately conducted. Here, we investigated the performance of three popular search engines (SEQUEST, Mascot, and MS Amanda) in conjunction with five filtering approaches, including respective score-based filtering, a group-based approach, local false discovery rate (LFDR), PeptideProphet, and Percolator. A total of eight data sets from various proteomes (e.g., E. coli, yeast, and human) produced by various instruments with high-accuracy survey scan (MS1) and high- or low-accuracy fragment ion scan (MS2) (LTQ-Orbitrap, Orbitrap-Velos, Orbitrap-Elite, Q-Exactive, Orbitrap-Fusion, and Q-TOF) were analyzed. It was found combinations involving Percolator achieved markedly more peptide and protein identifications at the same FDR level than the other 12 combinations for all data sets. Among these, combinations of SEQUEST–Percolator and MS Amanda–Percolator provided slightly better performances for data sets with low-accuracy MS2 (ion trap or IT) and high accuracy MS2 (Orbitrap or TOF), respectively, than did other methods. For approaches without Percolator, SEQUEST–group performs the best for data sets with MS2 produced by collision-induced dissociation (CID) and IT analysis; Mascot–LFDR gives more identifications for data sets generated by higher-energy collisional dissociation (HCD) and analyzed in Orbitrap (HCD–OT) and in Orbitrap Fusion (HCD–IT); MS Amanda–Group excels for the Q-TOF data set and the Orbitrap Velos HCD–OT data set. Therefore, if Percolator was not used, a specific combination should be applied for each type of data set. Moreover, a higher percentage of multiple-peptide proteins and lower variation of protein spectral counts were observed when analyzing technical replicates using Percolator-associated combinations; therefore, Percolator enhanced the reliability for both identification and quantification. The analyses were performed using the specific programs embedded in Proteome Discoverer, Scaffold, and an in-house algorithm (Build Summary). These results provide valuable guidelines for the optimal interpretation of proteomic results and the development of fit-for-purpose protocols under different situations.