BuildSummary: Using a Group-Based Approach To Improve the Sensitivity of Peptide/Protein Identification in Shotgun Proteomics

BuildSummary: Using a Group-Based Approach To Improve the Sensitivity of Peptide/Protein Identification in Shotgun Proteomics
复制标题

DOI:
10.1021/pr200194p
复制
发表时间:
2012-03-01
影响因子:
4.4
通讯作者:
Zeng, Rong
Zeng, Rong
中科院分区:
生物学2区
文献类型:
--
作者:
Sheng, Quanhu;Dai, Jie;Zeng, Rong

文献摘要

被引文献

相似文献

目标-诱饵数据库搜索策略是一种被广泛接受的估计多肽识别错误发现率(FDR)的标准方法,基于该方法可以从目标数据库中筛选出多肽-谱匹配(PSM)。为了在给定固定准确度(通常由蛋白质FDR阈值定义)的情况下提高蛋白质鉴定的灵敏度,通常使用后处理程序,该后处理程序集成了来自不同多肽搜索引擎的结果,所述不同的多肽搜索引擎已经分析了相同的数据集。在这项工作中,我们证明了根据前体电荷、缺失的内部切割位点的数量、修饰状态和蛋白酶末端的数量分组的PSM,以及按照其唯一的肽数分组的蛋白质应该根据给定的FDR单独过滤。我们还开发了一种迭代程序,根据给定的FDR同时过滤PSM和蛋白质。最后,我们提出了一个通用的框架来整合来自不同肽搜索引擎的结果,使用相同的FDR阈值。我们的方法是用多台LC/MS仪器从两个不同的生物样本中获取的几个鸟枪式蛋白质组数据集进行测试的。结果表明,该方法的性能令人满意。我们在一个用户友好的名为BuildSummary的软件包中实现了该方法,该软件包可以作为软件套件ProteomicsTools的一部分从http://www.proteomics.ac.cn/software/proteomicstools/index.htm免费下载。
The target-decoy database search strategy is widely accepted as a standard method for estimating the false discovery rate (FDR) of peptide identification, based on which peptide-spectrum matches (PSMs) from the target database are filtered. To improve the sensitivity of protein identification given a fixed accuracy (frequently defined by a protein FDR threshold), a postprocessing procedure is often used that integrates results from different peptide search engines that had assayed the same data set. In this work, we show that PSMs that are grouped by the,precursor charge, the number of missed internal cleavage sites, the modification state, and the numbers of protease termini and that the proteins grouped by their unique peptide count should be filtered separately according to the given FDR. We also develop an iterative procedure to filter the PSMs and proteins simultaneously, according to the given FDR. Finally, we present a general framework to integrate the results from different peptide search engines using the same FDR threshold. Our method was tested with several shotgun proteomics data sets that were acquired by multiple LC/MS instruments from two different biological samples. The results showed a satisfactory performance. We implemented the method in a user-friendly software package called BuildSummary, which can be downloaded for free from http://www.proteomics.ac.cn/software/proteomicstools/index.htm as part of the software suite ProteomicsTools.