Statistical calibration of the SEQUEST XCorr function.

Statistical calibration of the SEQUEST XCorr function.
复制标题

DOI:
10.1021/pr8011107
复制
发表时间:
2009-04
影响因子:
4.4
通讯作者:
Noble WS
Noble WS
中科院分区:
生物学2区
文献类型:
--
作者:
Klammer AA;Park CY;Noble WS

文献摘要

参考文献

被引文献

相似文献

从鸟枪法蛋白质组学液相色谱串联质谱 (LC-MS/MS) 实验中获得准确的肽鉴定需要一个评分函数,该函数始终将正确的肽谱匹配 (PSM) 排在错误匹配之上。我们观察到,对于 Sequest 得分函数 X corr,无法区分正确和不正确的 PSM 部分是由于得分分布的频谱特定属性。换句话说,一些光谱得分很高,无论针对哪些肽进行评分,而其他光谱得分很高,因为它们针对大​​量肽进行评分。我们描述了一种用于校准 PSM 评分函数的协议,并演示了其在 X corr 和初步 Sequest 评分函数 Sp 中的应用。该协议通过仅使用该谱图的分数分布单独计算每个谱图的 p 值来解释谱图和肽的特定效应。我们证明这些计算出的 p 值在零分布下是均匀的,因此可以准确测量显着性。这些 p 值可用于估计错误发现率,因此无需对诱饵数据库进行额外搜索。此外,我们表明 p 值比其基础分数更好地校准;因此,当对多个光谱中得分最高的 PSM 进行排名时,p 值可以更好地区分正确和不正确的 PSM。校准协议通常适用于可以识别适当参数族的任何 PSM 评分函数。
Obtaining accurate peptide identifications from shotgun proteomics liquid chromatography tandem mass spectrometry (LC-MS/MS) experiments requires a score function that consistently ranks correct peptide-spectrum matches (PSMs) above incorrect matches. We have observed that, for the Sequest score function X corr, the inability to discriminate between correct and incorrect PSMs is due in part to spectrum-specific properties of the score distribution. In other words, some spectra score well regardless of which peptides they are scored against, and other spectra score well because they are scored against a large number of peptides. We describe a protocol for calibrating PSM score functions, and we demonstrate its application to X corr and the preliminary Sequest score function Sp. The protocol accounts for spectrum- and peptide-specific effects by calculating p values for each spectrum individually, using only that spectrum’s score distribution. We demonstrate that these calculated p values are uniform under a null distribution and therefore accurately measure significance. These p values can be used to estimate the false discovery rate, therefore eliminating the need for an extra search against a decoy database. In addition, we show that the p values are better calibrated than their underlying scores; consequently, when ranking top-scoring PSMs from multiple spectra, p values are better at discriminating between correct and incorrect PSMs. The calibration protocol is generally applicable to any PSM score function for which an appopriate parametric family can be identified.
DOI: 10.1038/85686
发表时间: 2001-03-01
影响因子: 46.9
作者:
Washburn, MP;Wolters, D;Yates, JR
通讯作者: Yates, JR
DOI: 10.1021/pr0605320
发表时间: 2007-01-01
影响因子: 4.4
作者:
Higgs, Richard E.;Knierman, Michael D.;Hale, John E.
通讯作者: Hale, John E.
DOI: 10.1021/pr800127y
发表时间: 2008-07-01
影响因子: 4.4
作者:
Park, Christopher Y.;Klammer, Aaron A.;Noble, William S.
通讯作者: Noble, William S.
DOI: 10.1093/bioinformatics/bth092
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Craig, R;Beavis, RC
通讯作者: Beavis, RC
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y