Evaluation of linear models and missing value imputation for the analysis of peptide-centric proteomics

Evaluation of linear models and missing value imputation for the analysis of peptide-centric proteomics
复制标题

DOI:
10.1186/s12859-019-2619-6
复制
发表时间:
2019-03-01
期刊:
影响因子:
3
通讯作者:
Popescu, George V.
Popescu, George V.
中科院分区:
生物学4区
文献类型:
--
作者:
Berg, Philip;McConnell, Evan W.;Popescu, George V.

文献摘要

被引文献

相似文献

背景评价了通过液相色谱-质谱法处理自下而上蛋白质组学产生的数据的几种方法,特别是用于处理翻译后修饰(PTM)分析(如可逆半胱氨酸氧化)的以肽为中心的定量。本文提出了一种管道的基础上的R编程语言来分析PTM从肽为中心的无标记定量蛋白质组学data.ResultsOur方法,包括方差稳定化,归一化,缺失的数据填补占PTM测量的大动态范围。它还校正了富集方案的偏差,并减少了与无标记定量相关的随机和系统误差。通过使用线性模型分析(limma)进行全蛋白质组差异PTM定量来测试方法的性能。我们客观地比较两种插补方法沿着与显着性检验时,使用多重插补缺失data.ConclusionIdentifying PTM在大规模datasets是一个问题,需要新的方法来处理缺失数据插补和差异蛋白质组分析的独特的特点。线性模型结合多重插补可以显着优于基于t检验的决策方法。
BackgroundSeveral methods to handle data generated from bottom-up proteomics via liquid chromatography-mass spectrometry, particularly for peptide-centric quantification dealing with post-translational modification (PTM) analysis like reversible cysteine oxidation are evaluated. The paper proposes a pipeline based on the R programming language to analyze PTMs from peptide-centric label-free quantitative proteomics data.ResultsOur methodology includes variance stabilization, normalization, and missing data imputation to account for the large dynamic range of PTM measurements. It also corrects biases from an enrichment protocol and reduces the random and systematic errors associated with label-free quantification. The performance of the methodology is tested by performing proteome-wide differential PTM quantitation using linear models analysis (limma). We objectively compare two imputation methods along with significance testing when using multiple-imputation for missing data.ConclusionIdentifying PTMs in large-scale datasets is a problem with distinct characteristics that require new methods for handling missing data imputation and differential proteome analysis. Linear models in combination with multiple-imputation could significantly outperform a t-test-based decision method.