State-of-the art data normalization methods improve NMR-based metabolomic analysis.

State-of-the art data normalization methods improve NMR-based metabolomic analysis.
复制标题

DOI:
10.1007/s11306-011-0350-z
复制
发表时间:
2012-06
期刊:
影响因子:
3.6
通讯作者:
Gronwald, Wolfram
Gronwald, Wolfram
中科院分区:
医学3区
文献类型:
--
作者:
Kohl, Stefanie M.;Klein, Matthias S.;Hochrein, Jochen;Oefner, Peter J.;Spang, Rainer;Gronwald, Wolfram

文献摘要

参考文献

被引文献

相似文献

通过多元数据分析从大型代谢组学数据集中提取生物医学信息是相当复杂的。常见的挑战包括筛选差异产生的代谢物、估计倍数变化和样品分类等。在这些分析步骤之前,重要的是要尽量减少不必要的偏差和实验方差的贡献。这就是数据预处理的目的。在这项工作中,不同的数据归一化方法进行了系统的比较,采用两个不同的数据集产生的核磁共振(NMR)光谱。为此,使用了两种不同类型的归一化方法,一种旨在去除不需要的样品间变异,另一种通过变量缩放和方差稳定方法调整不同代谢物的方差。对从健康志愿者和患有常染色体多囊肾病(ADPKD)的患者获得的尿NMR指纹进行了评价,所有测试方法对样本分类的影响。根据拉丁方设计,在数据集上研究了筛选差异产生的代谢物的性能,其中将不同量的8种不同代谢物加标至人尿液基质中,同时保持总加标量恒定。此外,还进行了具体测试,系统地考察了不同预处理方法对分析数据结构的影响。总之,最初为DNA微阵列分析开发的预处理方法,特别是分位数和三次样条归一化,在减少偏倚、准确检测倍数变化和分类样品方面表现最好。本文的在线版本(doi:10.1007/s11306-011-0350-z)包含补充材料,可供授权用户使用。
Extracting biomedical information from large metabolomic datasets by multivariate data analysis is of considerable complexity. Common challenges include among others screening for differentially produced metabolites, estimation of fold changes, and sample classification. Prior to these analysis steps, it is important to minimize contributions from unwanted biases and experimental variance. This is the goal of data preprocessing. In this work, different data normalization methods were compared systematically employing two different datasets generated by means of nuclear magnetic resonance (NMR) spectroscopy. To this end, two different types of normalization methods were used, one aiming to remove unwanted sample-to-sample variation while the other adjusts the variance of the different metabolites by variable scaling and variance stabilization methods. The impact of all methods tested on sample classification was evaluated on urinary NMR fingerprints obtained from healthy volunteers and patients suffering from autosomal polycystic kidney disease (ADPKD). Performance in terms of screening for differentially produced metabolites was investigated on a dataset following a Latin-square design, where varied amounts of 8 different metabolites were spiked into a human urine matrix while keeping the total spike-in amount constant. In addition, specific tests were conducted to systematically investigate the influence of the different preprocessing methods on the structure of the analyzed data. In conclusion, preprocessing methods originally developed for DNA microarray analysis, in particular, Quantile and Cubic-Spline Normalization, performed best in reducing bias, accurately detecting fold changes, and classifying samples. The online version of this article (doi:10.1007/s11306-011-0350-z) contains supplementary material, which is available to authorized users.
DOI: 10.1021/pr101080e
发表时间: 2011-03-01
影响因子: 4.4
作者:
Keeping, Andrew J.;Collins, Richard A.
通讯作者: Collins, Richard A.
DOI: 10.1177/0192623307310947
发表时间: 2008-01-01
影响因子: 1.5
作者:
Clarke, Christopher J.;Haselden, John N.
通讯作者: Haselden, John N.
DOI: 10.2307/2987937
发表时间: 1983-01-01
期刊: JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES D-THE STATISTICIAN
影响因子: --
作者:
ALTMAN, DG;BLAND, JM
通讯作者: BLAND, JM
DOI: 10.1007/978-1-60761-987-1_16
发表时间: 2011-01-01
期刊: DATA MINING IN PROTEOMICS: FROM STANDARDS TO APPLICATIONS
影响因子: --
作者:
Jung, Klaus
通讯作者: Jung, Klaus
DOI: 10.2307/2289282
发表时间: 1988-09-01
影响因子: 3.7
作者:
CLEVELAND, WS;DEVLIN, SJ
通讯作者: DEVLIN, SJ