State-of-the art data normalization methods improve NMR-based metabolomic analysis.
State-of-the art data normalization methods improve NMR-based metabolomic analysis.
复制标题
DOI:
10.1007/s11306-011-0350-z
复制
发表时间:
2012-06
期刊:
影响因子:
3.6
通讯作者:
Gronwald, Wolfram
中科院分区:
文献类型:
--
作者:
Kohl, Stefanie M.;Klein, Matthias S.;Hochrein, Jochen;Oefner, Peter J.;Spang, Rainer;Gronwald, Wolfram
Extracting biomedical information from large metabolomic datasets by multivariate data analysis is of considerable complexity. Common challenges include among others screening for differentially produced metabolites, estimation of fold changes, and sample classification. Prior to these analysis steps, it is important to minimize contributions from unwanted biases and experimental variance. This is the goal of data preprocessing. In this work, different data normalization methods were compared systematically employing two different datasets generated by means of nuclear magnetic resonance (NMR) spectroscopy. To this end, two different types of normalization methods were used, one aiming to remove unwanted sample-to-sample variation while the other adjusts the variance of the different metabolites by variable scaling and variance stabilization methods. The impact of all methods tested on sample classification was evaluated on urinary NMR fingerprints obtained from healthy volunteers and patients suffering from autosomal polycystic kidney disease (ADPKD). Performance in terms of screening for differentially produced metabolites was investigated on a dataset following a Latin-square design, where varied amounts of 8 different metabolites were spiked into a human urine matrix while keeping the total spike-in amount constant. In addition, specific tests were conducted to systematically investigate the influence of the different preprocessing methods on the structure of the analyzed data. In conclusion, preprocessing methods originally developed for DNA microarray analysis, in particular, Quantile and Cubic-Spline Normalization, performed best in reducing bias, accurately detecting fold changes, and classifying samples. The online version of this article (doi:10.1007/s11306-011-0350-z) contains supplementary material, which is available to authorized users.
登录
查看更多内容
影响因子:
4.4
作者:
Keeping, Andrew J.;Collins, Richard A.
通讯作者:
Collins, Richard A.
影响因子:
1.5
作者:
Clarke, Christopher J.;Haselden, John N.
通讯作者:
Haselden, John N.
DOI:
10.2307/2987937
发表时间:
1983-01-01
期刊:
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES D-THE STATISTICIAN
影响因子:
--
作者:
ALTMAN, DG;BLAND, JM
通讯作者:
BLAND, JM
DOI:
10.1007/978-1-60761-987-1_16
发表时间:
2011-01-01
期刊:
DATA MINING IN PROTEOMICS: FROM STANDARDS TO APPLICATIONS
影响因子:
--
作者:
Jung, Klaus
通讯作者:
Jung, Klaus
影响因子:
3.7
作者:
CLEVELAND, WS;DEVLIN, SJ
通讯作者:
DEVLIN, SJ