A cross-validation scheme for machine learning algorithms in shotgun proteomics.

A cross-validation scheme for machine learning algorithms in shotgun proteomics.
复制标题

DOI:
10.1186/1471-2105-13-s16-s3
复制
发表时间:
2012
期刊:
影响因子:
3
通讯作者:
Käll L
Käll L
中科院分区:
生物学4区
文献类型:
--
作者:
Granholm V;Noble WS;Käll L

文献摘要

被引文献

相似文献

通过将观察到的光谱与来自蛋白质数据库的肽相匹配,通常可以从基于质谱的蛋白质组学实验中鉴定肽。这些识别的错误率可以通过目标-诱饵分析来估计,这涉及到将光谱与洗牌或逆转的肽相匹配。除了估计错误率外,半监督机器学习算法还可以使用诱饵搜索来增加自信识别的肽的数量。然而,对于所有的机器学习算法,必须对结果进行验证,以避免过度拟合或有偏差的学习等问题,这些问题会产生不可靠的肽识别。在这里,我们讨论了如何将目标诱饵方法应用于霰弹枪蛋白质组学的机器学习中,重点讨论了如何通过交叉验证(机器学习中常用的验证方案)来验证结果。我们还使用模拟数据来证明所提出的交叉验证方案检测过拟合的能力。
Peptides are routinely identified from mass spectrometry-based proteomics experiments by matching observed spectra to peptides derived from protein databases. The error rates of these identifications can be estimated by target-decoy analysis, which involves matching spectra to shuffled or reversed peptides. Besides estimating error rates, decoy searches can be used by semi-supervised machine learning algorithms to increase the number of confidently identified peptides. As for all machine learning algorithms, however, the results must be validated to avoid issues such as overfitting or biased learning, which would produce unreliable peptide identifications. Here, we discuss how the target-decoy method is employed in machine learning for shotgun proteomics, focusing on how the results can be validated by cross-validation, a frequently used validation scheme in machine learning. We also use simulated data to demonstrate the proposed cross-validation scheme's ability to detect overfitting.