Improving sensitivity in proteome studies by analysis of false discovery rates for multiple search engines.

Improving sensitivity in proteome studies by analysis of false discovery rates for multiple search engines.
复制标题

DOI:
10.1002/pmic.200800473
复制
发表时间:
2009-03
期刊:
影响因子:
3.4
通讯作者:
Paton, Norman W.
Paton, Norman W.
中科院分区:
生物学3区
文献类型:
--
作者:
Jones, Andrew R.;Siepen, Jennifer A.;Hubbard, Simon J.;Paton, Norman W.

文献摘要

参考文献

被引文献

相似文献

Tandem mass spectrometry, run in combination with liquid chromatography (LC-MS/MS), can generate large numbers of peptide and protein identifications, for which a variety of database search engines are available. Distinguishing correct identifications from false positives is far from trivial because all data sets are noisy, and tend to be too large for manual inspection, therefore probabilistic methods must be employed to balance the trade-off between sensitivity and specificity. Decoy databases are becoming widely used to place statistical confidence in results sets, allowing the false discovery rate (FDR) to be estimated. It has previously been demonstrated that different MS search engines produce different peptide identification sets, and as such, employing more than one search engine could result in an increased number of peptides being identified. However, such efforts are hindered by the lack of a single scoring framework employed by all search engines. We have developed a search engine independent scoring framework based on FDR which allows peptide identifications from different search engines to be combined, called the FDRScore. We observe that peptide identifications made by three search engines are infrequently false positives, and identifications made by only a single search engine, even with a strong score from the source search engine, are significantly more likely to be false positives. We have developed a second score based on the FDR within peptide identifications grouped according to the set of search engines that have made the identification, called the combined FDRScore. We demonstrate by searching large publicly available data sets that the combined FDRScore can differentiate between between correct and incorrect peptide identifications with high accuracy, allowing on average 35% more peptide identifications to be made at a fixed FDR than using a single search engine.
DOI: 10.1021/ac025826t
发表时间: 2002-11-01
影响因子: 7.4
作者:
MacCoss, MJ;Wu, CC;Yates, JR
通讯作者: Yates, JR
DOI: 10.1021/ac0258709
发表时间: 2003-02-15
影响因子: 7.4
作者:
Fenyö, D;Beavis, RC
通讯作者: Beavis, RC
DOI: 10.1002/pmic.200300485
发表时间: 2003-08-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Colinge, J;Masselot, A;Magnin, J
通讯作者: Magnin, J
DOI: 10.1021/ac0341261
发表时间: 2003-09-01
影响因子: 7.4
作者:
Nesvizhskii, AI;Keller, A;Aebersold, R
通讯作者: Aebersold, R
PeptideAtlas项目。
DOI: 10.1093/nar/gkj040
发表时间: 2006-01-01
影响因子: 14.9
作者:
通讯作者: --