Combined Statistical Analyses of Peptide Intensities and Peptide Occurrences Improves Identification of Significant Peptides from MS-Based Proteomics Data

Combined Statistical Analyses of Peptide Intensities and Peptide Occurrences Improves Identification of Significant Peptides from MS-Based Proteomics Data
复制标题

DOI:
10.1021/pr1005247
复制
发表时间:
2010-11-01
影响因子:
4.4
通讯作者:
Pounds, Joel G.
Pounds, Joel G.
中科院分区:
生物学2区
文献类型:
--
作者:
Webb-Robertson, Bobbie-Jo M.;McCue, Lee Ann;Pounds, Joel G.

文献摘要

被引文献

相似文献

基于液相色谱-质谱法 (LC-MS) 的蛋白质组学利用蛋白水解肽的峰值强度来推断肽/蛋白质的差异丰度。然而,肽的强度和观察(存在/不存在)的运行间的巨大差异使得数据分析变得非常具有挑战性。 LC-MS 蛋白质组学数据中缺失的观察结果很难用传统的基于插补的方法来解决,因为数据缺失的机制是先验未知的。由于实验误差等随机机制或真实生物效应等非随机机制,数据可能会丢失。我们提出了一种统计方法,使用称为 G 检验的独立性检验来检验实验组中缺失值数量之间的独立性零假设。我们将 G 检验结果配对,通过方差分析 (ANOVA) 评估缺失数据 (IMD) 的独立性,方差分析 (ANOVA) 仅使用从观测数据计算得出的均值和方差。因此,每个肽都由两种统计置信度指标表示,一种用于定性差异观察,另一种用于定量差异强度。我们使用三个 LC-MS 数据集来证明 IMD-ANOVA 方法的稳健性和灵敏度。
Liquid chromatography-mass spectrometry-based (LC-MS) proteomics uses peak intensities of proteolytic peptides to infer the differential abundance of peptides/proteins. However, substantial run-to-run variability in intensities and observations (presence/absence) of peptides makes data analysis quite challenging. The missing observations in LC-MS proteomics data are difficult to address with traditional imputation-based approaches because the mechanisms by which data are missing are unknown a priori. Data can be missing due to random mechanisms such as experimental error or nonrandom mechanisms such as a true biological effect. We present a statistical approach that uses a test of independence known as a G-test to test the null hypothesis of independence between the number of missing values across experimental groups. We pair the G-test results, evaluating independence of missing data (IMD) with an analysis of variance (ANOVA) that uses only means and variances computed from the observed data. Each peptide is therefore represented by two statistical confidence metrics, one for qualitative differential observation and one for quantitative differential intensity. We use three LC-MS data sets to demonstrate the robustness and sensitivity of the IMD-ANOVA approach.