A survey of computational methods and error rate estimation procedures for peptide and protein identification in shotgun proteomics.

A survey of computational methods and error rate estimation procedures for peptide and protein identification in shotgun proteomics.
复制标题

DOI:
10.1016/j.jprot.2010.08.009
复制
发表时间:
2010-10-10
影响因子:
3.3
通讯作者:
Nesvizhskii AI
Nesvizhskii AI
中科院分区:
生物学2区
文献类型:
--
作者:
Nesvizhskii AI

文献摘要

参考文献

被引文献

相似文献

这份手稿提供了一个全面的审查肽和蛋白质鉴定过程中使用串联质谱(MS/MS)数据产生的鸟枪蛋白质组学实验。常用的方法分配肽序列的MS/MS光谱进行了严格的讨论和比较,从基本的策略,先进的多阶段的方法。特别注意的是假阳性鉴定的问题。现有的统计方法用于评估肽的重要性,光谱匹配进行了调查,范围从单光谱的方法,如期望值的全球错误率估计程序,如错误发现率和后验概率。使用辅助判别信息(质量准确度、肽分离坐标、消化特性等)的重要性的讨论,并提出了先进的计算方法,联合建模的多个信息源。这篇综述还包括对影响蛋白质水平数据解释的问题的详细分析,包括从肽到蛋白质水平时错误率的放大,以及在存在共享肽的情况下推断样品蛋白质身份的模糊性。常用的方法计算蛋白质水平的置信分数进行了详细讨论。审查的结论与几个突出的计算问题的讨论。
This manuscript provides a comprehensive review of the peptide and protein identification process using tandem mass spectrometry (MS/MS) data generated in shotgun proteomic experiments. The commonly used methods for assigning peptide sequences to MS/MS spectra are critically discussed and compared, from basic strategies to advanced multi-stage approaches. A particular attention is paid to the problem of false-positive identifications. Existing statistical approaches for assessing the significance of peptide to spectrum matches are surveyed, ranging from single-spectrum approaches such as expectation values to global error rate estimation procedures such as false discovery rates and posterior probabilities. The importance of using auxiliary discriminant information (mass accuracy, peptide separation coordinates, digestion properties, and etc.) is discussed, and advanced computational approaches for joint modeling of multiple sources of information are presented. This review also includes a detailed analysis of the issues affecting the interpretation of data at the protein level, including the amplification of error rates when going from peptide to protein level, and the ambiguities in inferring the identifies of sample proteins in the presence of shared peptides. Commonly used methods for computing protein-level confidence scores are discussed in detail. The review concludes with a discussion of several outstanding computational issues.
DOI: 10.1371/journal.pone.0008949
发表时间: 2010-01-28
期刊: PloS one
影响因子: 3.7
作者:
Bitton DA;Smith DL;Connolly Y;Scutt PJ;Miller CJ
通讯作者: Miller CJ
DOI: 10.1021/pr800917p
发表时间: 2009-04-01
影响因子: 4.4
作者:
Bailey, Christopher M.;Sweet, Steve M. M.;Cooper, Helen J.
通讯作者: Cooper, Helen J.
DOI: 10.1186/1745-6150-3-27
发表时间: 2008-07-02
期刊: BIOLOGY DIRECT
影响因子: 5.5
作者:
Alves, Gelio;Ogurtsov, Aleksey Y.;Yu, Yi-Kuo
通讯作者: Yu, Yi-Kuo
DOI: 10.1002/elps.200900332
发表时间: 2009-11-01
期刊: ELECTROPHORESIS
影响因子: 2.9
作者:
Bertsch, Andreas;Leinenbach, Andreas;Kohlbacher, Oliver
通讯作者: Kohlbacher, Oliver
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y