Estimating the statistical significance of peptide identifications from shotgun proteomics experiments

Estimating the statistical significance of peptide identifications from shotgun proteomics experiments
复制标题

DOI:
10.1021/pr0605320
复制
发表时间:
2007-01-01
影响因子:
4.4
通讯作者:
Hale, John E.
Hale, John E.
中科院分区:
生物学2区
文献类型:
--
作者:
Higgs, Richard E.;Knierman, Michael D.;Hale, John E.

文献摘要

被引文献

相似文献

我们提出了一种基于包装器的方法,使用来自多个商业MS/MS搜索引擎的输出来估计和控制多肽鉴定的错误发现率。该方法的特征包括在灵活的分类模型中将来自多个搜索引擎的输出与序列和光谱派生特征相结合的灵活性,以产生与正确的肽识别相关联的分数。来自反向数据库搜索的该分类模型分数被用作零分布,用于使用简单且已建立的统计过程来估计p值和错误发现率。在LTQ-FT质谱仪上对大鼠血清的10个分析结果表明,该方法在控制一组报告的多肽鉴定中的假阳性比例方面得到了很好的校准,同时正确识别了比仅使用一个搜索引擎的基于规则的方法更多的多肽。关键词:多肽鉴定中心点假发现率中心点SEQUEST中心点X!串联中心点统计意义中心点蛋白质组学
We present a wrapper-based approach to estimate and control the false discovery rate for peptide identifications using the outputs from multiple commercially available MS/MS search engines. Features of the approach include the flexibility to combine output from multiple search engines with sequence and spectral derived features in a flexible classification model to produce a score associated with correct peptide identifications. This classification model score from a reversed database search is taken as the null distribution for estimating p-values and false discovery rates using a simple and established statistical procedure. Results from 10 analyses of rat sera on an LTQ-FT mass spectrometer indicate that the method is well calibrated for controlling the proportion of false positives in a set of reported peptide identifications while correctly identifying more peptides than rule-based methods using one search engine alone.Keywords: peptide identification center dot false discovery rate center dot Sequest center dot X! Tandem center dot statistical significance center dot proteomics