Normalization and Statistical Analysis of Quantitative Proteomics Data Generated by Metabolic Labeling

Normalization and Statistical Analysis of Quantitative Proteomics Data Generated by Metabolic Labeling
复制标题

DOI:
10.1074/mcp.m800462-mcp200
复制
发表时间:
2009-10-01
影响因子:
7
通讯作者:
Cavicchioli, Ricardo
Cavicchioli, Ricardo
中科院分区:
生物学1区
文献类型:
--
作者:
Ting, Lily;Cowley, Mark J.;Cavicchioli, Ricardo

文献摘要

被引文献

相似文献

比较蛋白质组学是了解生物系统对生长参数变化的反应的一种强大的分析方法。为了对生物反应做出可靠的推断,蛋白质组学方法必须结合定量数据的适当统计措施。在本工作中,我们应用基于微阵列的归一化和统计分析(显著性检验)方法来分析一种海洋细菌(Sphingopyxis alaskensis)代谢标记产生的定量蛋白质组学数据。生成了1172个蛋白质的定量数据,代表了1736个高置信度蛋白质鉴定(54%的基因组覆盖率)。为了测试归一化的方法,细胞在单一温度下生长,用N-14或N-15进行代谢标记,并以不同的比例组合,以获得人为扭曲的数据集。比率与平均值(MA)图的检验确定了固定值的中位数归一化最适合于数据。为了确定评估差异丰度的适当统计方法,对两种温度下生长的细胞的蛋白质组学数据应用了a倍变化法、Student’st检验、无调节t检验和经验贝叶斯调节t检验。通过多次技术和生物复制使用反向代谢标记,并对基于相同培养光密度组合的细胞(提供偏斜数据)或组合以提供等量蛋白质(无偏斜)的细胞提取物进行蛋白质组学。为了考虑任意复杂的实验特定参数,使用R/Bioconductor中的limma软件包使用线性建模方法来分析数据。通过使用低归一化(在MA图检验后)和应用经验贝叶斯缓和t检验,获得了具有统计显著性差异丰富蛋白的高质量列表。该方法还有效地控制了错误发现的数量,并使用story -Tibshirani错误发现率(Storey, J. D., and Tibshirani, R.(2003)全基因组研究的统计显著性来纠正多重测试问题。Proc。国家的。学会科学。《美国法典》100,94409445)。我们开发的方法一般适用于多种生物系统的定量蛋白质组学分析。中国生物医学工程学报,2009;
Comparative proteomics is a powerful analytical method for learning about the responses of biological systems to changes in growth parameters. To make confident inferences about biological responses, proteomics approaches must incorporate appropriate statistical measures of quantitative data. In the present work we applied microarray-based normalization and statistical analysis (significance testing) methods to analyze quantitative proteomics data generated from the metabolic labeling of a marine bacterium (Sphingopyxis alaskensis). Quantitative data were generated for 1,172 proteins, representing 1,736 high confidence protein identifications (54% genome coverage). To test approaches for normalization, cells were grown at a single temperature, metabolically labeled with N-14 or N-15, and combined in different ratios to give an artificially skewed data set. Inspection of ratio versus average (MA) plots determined that a fixed value median normalization was most suitable for the data. To determine an appropriate statistical method for assessing differential abundance, a-fold change approach, Student's t test, unmoderated t test, and empirical Bayes moderated t test were applied to proteomics data from cells grown at two temperatures. Inverse metabolic labeling was used with multiple technical and biological replicates, and proteomics was performed on cells that were combined based on equal optical density of cultures (providing skewed data) or on cell extracts that were combined to give equal amounts of protein (no skew). To account for arbitrarily complex experiment-specific parameters, a linear modeling approach was used to analyze the data using the limma package in R/Bioconductor. A high quality list of statistically significant differentially abundant proteins was obtained by using lowess normalization (after inspection of MA plots) and applying the empirical Bayes moderated t test. The approach also effectively controlled for the number of false discoveries and corrected for the multiple testing problem using the Storey-Tibshirani false discovery rate (Storey, J. D., and Tibshirani, R. (2003) Statistical significance for genome-wide studies. Proc. Natl. Acad. Sci. U. S. A. 100, 94409445). The approach we have developed is generally applicable to quantitative proteomics analyses of diverse biological systems. Molecular & Cellular Proteomics 8: 2227-2242, 2009.