Analysis of the Resolution Limitations of Peptide Identification Algorithms

Analysis of the Resolution Limitations of Peptide Identification Algorithms
复制标题

DOI:
10.1021/pr200913a
复制
发表时间:
2011-12-01
影响因子:
4.4
通讯作者:
Martens, Lennart
Martens, Lennart
中科院分区:
生物学2区
文献类型:
--
作者:
Colaert, Niklaas;Degroeve, Sven;Martens, Lennart

文献摘要

被引文献

相似文献

使用肽中心蛋白质组学技术的蛋白质组鉴定是常规使用的分析技术。从MS/MS谱图中识别肽的最强大和最流行的方法之一是使用搜索引擎的蛋白质数据库匹配。通过靶/诱饵搜索的假发现率(FDR)估计的显著性阈值用于确保保留MS/MS谱对肽的主要置信度分配。然而,当使用这种诱饵搜索来估计FDR时,缺点已经变得明显。为了研究这些缺点,我们在这里介绍了一种新的诱饵数据库,其中包含在原始搜索中识别的肽的同量异位素突变版本。由于产生诱捕序列的监督方式,我们称之为定向诱饵数据库。由于在我们的定向诱饵数据库中发现的肽因此被专门设计为看起来与正向识别非常相似,因此可以分析现有搜索算法在这种强烈混淆的情况下进行正确调用的局限性。有趣的是,对于绝大多数确定的肽鉴定,可以发现定向诱饵肽与谱匹配具有比正向匹配分数更好或相等的匹配分数,突出了当今高通量蛋白质组学中肽鉴定解释的重要问题。
Proteome identification using peptide-centric proteomics techniques is a routinely used analysis technique. One of the most powerful and popular methods for the identification of peptides from MS/MS spectra is protein database matching using search engines. Significance thresholding through false discovery rate (FDR) estimation by target/decoy searches is used to ensure the retention of predominantly confident assignments of MS/MS spectra to peptides. However, shortcomings have become apparent when such decoy searches are used to estimate the FDR. To study these shortcomings, we here introduce a novel kind of decoy database that contains isobaric mutated versions of the peptides that were identified in the original search. Because of the supervised way in which the entrapment sequences are generated, we call this a directed decoy database. Since the peptides found in our directed decoy database are thus specifically designed to look quite similar to the forward identifications, the limitations of the existing search algorithms in making correct calls in such strongly confusing situations can be analyzed. Interestingly, for the vast majority of confidently identified peptide identifications, a directed decoy peptide-to-spectrum match can be found that has a better or equal match score than the forward match score, highlighting an important issue in the interpretation of peptide identifications in present-day high-throughput proteomics.