A predictive model for identifying proteins by a single peptide match

A predictive model for identifying proteins by a single peptide match
复制标题

DOI:
10.1093/bioinformatics/btl595
复制
发表时间:
2007-02-01
期刊:
影响因子:
5.8
通讯作者:
Kolker, Eugene
Kolker, Eugene
中科院分区:
生物学3区
文献类型:
--
作者:
Higdon, Roger;Kolker, Eugene

文献摘要

被引文献

相似文献

动机:胰酶消化的串联质谱学,然后是数据库搜索,是高通量蛋白质组学研究中最受欢迎的方法之一。如果肽超过一定的评分阈值,则被认为是已识别的。为避免假阳性蛋白质鉴定,通常推荐在单一蛋白质内鉴定2个唯一的多肽。尽管如此,在一个典型的高通量实验中,数百种蛋白质仅由一种多肽识别。在这里,我们介绍了一种方法来区分真假识别之间的一次击中的蛋白质。该方法基于随机数据库搜索和使用具有交叉验证的Logistic回归模型。该方法被应用于三个细菌样本的分析,能够回收68-98%的正确的单次命中蛋白质,错误率为2%。这导致识别的蛋白质数量增加了22%-65%。鉴定真正的单次敲击蛋白质将导致发现许多关键的调节因子、生物标志物和其他低丰度蛋白质。
Motivation: Tandem mass-spectrometry of trypsin digests, followed by database searching, is one of the most popular approaches in high-throughput proteomics studies. Peptides are considered identified if they pass certain scoring thresholds. To avoid false positive protein identification, >= 2 unique peptides identified within a single protein are generally recommended. Still, in a typical high-throughput experiment, hundreds of proteins are identified only by a single peptide. We introduce here a method for distinguishing between true and false identifications among single-hit proteins. The approach is based on randomized database searching and usage of logistic regression models with cross-validation. This approach is implemented to analyze three bacterial samples enabling recovery 68-98% of the correct single-hit proteins with an error rate of < 2%. This results in a 22-65% increase in number of identified proteins. Identifying true single-hit proteins will lead to discovering many crucial regulators, biomarkers and other low abundance proteins.