Summarization vs Peptide-Based Models in Label-Free Quantitative Proteomics: Performance, Pitfalls, and Data Analysis Guidelines

Summarization vs Peptide-Based Models in Label-Free Quantitative Proteomics: Performance, Pitfalls, and Data Analysis Guidelines
复制标题

DOI:
10.1021/pr501223t
复制
发表时间:
2015-06-01
影响因子:
4.4
通讯作者:
Clement, Lieven
Clement, Lieven
中科院分区:
生物学2区
文献类型:
--
作者:
Goeminne, Ludger J. E.;Argentini, Andrea;Clement, Lieven

文献摘要

被引文献

相似文献

定量无标记质谱越来越多地用于分析复杂生物样品的蛋白质组。然而,选择适当的数据分析方法仍然是一个重大挑战。因此,我们提供了基于肽的模型和基于肽总结的管道之间的严格比较。我们表明,基于肽的模型在灵敏度,特异性,准确性和精确度方面优于基于汇总的管道。我们还表明,预定义的FDR截止值的差异调节蛋白的检测可能成为问题时,差异表达(DE)蛋白在一个或多个样品中是高度丰富的。因此,在解释加标内部质控品的样品和含有少量非常高丰度蛋白质的样品的数据时应小心。然而,我们确实表明,特定的诊断图可用于评估差异表达的蛋白质和获得的倍数变化估计的总体质量。最后,我们的研究还表明,“低丰度缺失”假设下的插补有利于检测低丰度蛋白质的差异表达,但它会对中丰度至高丰度蛋白质产生负面影响。因此,应谨慎使用标准蛋白质组学软件中常用的插补策略。
Quantitative label-free mass spectrometry is increasingly used to analyze the proteomes of complex biological samples. However, the choice of appropriate data analysis methods remains a major challenge. We therefore provide a rigorous comparison between peptide-based models and peptide-summarization-based pipelines. We show that peptide-based models outperform summarization-based pipelines in terms of sensitivity, specificity, accuracy, and precision. We also demonstrate that the predefined FDR cutoffs for the detection of differentially regulated proteins can become problematic when differentially expressed (DE) proteins are highly abundant in one or more samples. Care should therefore be taken when data are interpreted from samples with spiked-in internal controls and from samples that contain a few very highly abundant proteins. We do, however, show that specific diagnostic plots can be used for assessing differentially expressed proteins and the overall quality of the obtained fold change estimates. Finally, our study also illustrates that imputation under the "missing by low abundance" assumption is beneficial for the detection of differential expression in proteins with low abundance, but it negatively affects moderately to highly abundant proteins. Hence, imputation strategies that are commonly implemented in standard proteomics software should be used with care.