Improvements to the percolator algorithm for Peptide identification from shotgun proteomics data sets.

Improvements to the percolator algorithm for Peptide identification from shotgun proteomics data sets.
复制标题

DOI:
10.1021/pr801109k
复制
发表时间:
2009-07
影响因子:
4.4
通讯作者:
Noble WS
Noble WS
中科院分区:
生物学2区
文献类型:
--
作者:
Spivak M;Weston J;Bottou L;Käll L;Noble WS

文献摘要

参考文献

被引文献

相似文献

霰弹枪蛋白质组学与数据库搜索软件相结合,可以在单个实验中识别大量肽。然而,一些现有的搜索算法,如SEQUEST,使用评分函数,主要用于识别给定谱的最佳肽。因此,当跨光谱比较鉴定时,SEQUEST评分函数Xcorr无法准确区分正确和不正确的肽鉴定。已经提出了几种机器学习方法来解决区分正确和不正确肽谱匹配(psm)的分类任务。最近的一个例子是Percolator,它使用半监督学习和诱饵数据库搜索策略来学习区分由数据库搜索算法识别的正确和不正确的psm。目前的工作描述了对Percolator的三个改进。(1) Percolator的启发式优化被清晰的目标函数所取代,其选择背后有直观的原因。(2)使用可处理的非线性模型代替线性模型,使得精度比原来的Percolator提高。(3)提出了一种q值下直接优化识别光谱数量的方法q -ranker,取得了进一步的效果。
Shotgun proteomics coupled with database search software allows the identification of a large number of peptides in a single experiment. However, some existing search algorithms, such as SEQUEST, use score functions that are designed primarily to identify the best peptide for a given spectrum. Consequently, when comparing identifications across spectra, the SEQUEST score function Xcorr fails to discriminate accurately between correct and incorrect peptide identifications. Several machine learning methods have been proposed to address the resulting classification task of distinguishing between correct and incorrect peptide-spectrum matches (PSMs). A recent example is Percolator, which uses semi-supervised learning and a decoy database search strategy to learn to distinguish between correct and incorrect PSMs identified by a database search algorithm. The current work describes three improvements to Percolator. (1) Percolator’s heuristic optimization is replaced with a clear objective function, with intuitive reasons behind its choice. (2) Tractable nonlinear models are used instead of linear models, leading to improved accuracy over the original Percolator. (3) A method, Q-ranker, for directly optimizing the number of identified spectra at a specified q value is proposed, which achieves further gains.
DOI: 10.1093/bioinformatics/btn189
发表时间: 2008-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Klammer AA;Reynolds SM;Bilmes JA;MacCoss MJ;Noble WS
通讯作者: Noble WS
DOI: 10.1021/pr050315j
发表时间: 2006-03-01
影响因子: 4.4
作者:
Klammer, AA;MacCoss, MJ
通讯作者: MacCoss, MJ
DOI: 10.1111/1467-9868.00346
发表时间: 2002-01-01
影响因子: 5.8
作者:
Storey, JD
通讯作者: Storey, JD
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1007/bf00994018
发表时间: 1995-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
CORTES, C;VAPNIK, V
通讯作者: VAPNIK, V