Data Imputation in Merged Isobaric Labeling-Based Relative Quantification Datasets

Data Imputation in Merged Isobaric Labeling-Based Relative Quantification Datasets
复制标题

DOI:
10.1007/978-1-4939-9744-2_13
复制
发表时间:
2020-01-01
期刊:
MASS SPECTROMETRY DATA ANALYSIS IN PROTEOMICS, 3RD EDITION
影响因子:
--
通讯作者:
Beck, Hans Christian
Beck, Hans Christian
中科院分区:
其他
文献类型:
--
作者:
Palstrom, Nicolai Bjodstrup;Matthiesen, Rune;Beck, Hans Christian

文献摘要

被引文献

相似文献

基于质谱的蛋白质组学数据获取与同量异位素标记定量分析相结合(iTRAQ和TMT)不可避免地在蛋白质组学实验中引入缺失值,其中结合了许多LC运行,特别是在日益增长的鸟枪临床蛋白质组学领域,其中将来自几百个患者样品的蛋白质组学分析的蛋白质谱与临床特征(例如疾病或疾病治疗,以便将特定结果与一种或多种蛋白质联系起来。在临床研究的背景下,很明显,这些数据集中的缺失值降低了下游统计分析的能力,因此可能妨碍疾病性状的表达与可能用于预后、诊断或预测目的的特定蛋白质的表达的联系。在我们的研究中,我们测试了最初为微阵列数据开发的三种数据填补方法,用于填补由几次鸟枪蛋白质组学实验产生的数据集中的缺失值,其中数据是基于同量异位素标签(iTRAQ和TMT)的相对蛋白质丰度。我们的结论是,填补方法的基础上,k最近邻成功地填补缺失值的数据集高达50%的缺失值。
The data-dependent acquisition in mass spectrometry-based proteomics combined with quantitative analysis using isobaric labeling (iTRAQ and TMT) inevitably introduces missing values in proteomic experiments where a number of LC-runs are combined, especially in the growing field of shotgun clinical proteomics, where the protein profiles from the proteomics analysis of several hundred patient samples are compared and correlated to clinical traits such as a specific disease or disease treatment in order to link specific outcomes to one or more proteins. In the context of clinical research it is evident that missing values in such datasets reduce the power of the downstream statistical analysis therefore may hampers the linking of the expression of disease traits to the expression of specific proteins that may be useful for prognostic, diagnostic, or predictive purposes. In our study, we tested three data imputation approaches initially developed for microarray data for the imputation of missing values in datasets that are generated by several runs of shotgun proteomic experiments and where the data were relative protein abundances based on isobaric tags (iTRAQ and TMT). Our conclusion is that imputation methods based on k Nearest Neighbors successfully impute missing values in datasets with up to 50% missing values.