Comparing normalization methods and the impact of noise

Comparing normalization methods and the impact of noise
复制标题

DOI:
10.1007/s11306-018-1400-6
复制
发表时间:
2018-08-01
期刊:
影响因子:
3.6
通讯作者:
Powers, Robert
Powers, Robert
中科院分区:
医学3区
文献类型:
--
作者:
Thao Vu;Riekeberg, Eli;Powers, Robert

文献摘要

被引文献

相似文献

未能正确解释OMICS数据集中的正常系统变化可能会导致误导性的生物学结论。因此,标准化是OMICS数据集正确预处理的必要步骤。在这方面,最佳的归一化方法将有效地减少不必要的偏差,并提高下游定量分析的准确性。但是,目前还不清楚哪种归一化方法是最好的,因为每种算法以不同的方式处理系统噪声。确定代谢组学数据集预处理的归一化方法的最佳选择。九种MVAPACK归一化算法与添加高斯噪声和随机稀释因子修改的模拟和实验NMR谱进行了比较。方法进行了评估的基础上恢复的能力,真正的光谱峰的强度和真实的分类功能的再现性从正交投影到潜在结构判别分析模型(OPLS-DA)。大多数归一化方法(直方图匹配除外)在中等水平的信号方差同样表现良好。只有概率商(PQ)和常数和(CS)在最大噪声下保持了最高水平的峰恢复率(> 67%)和与真实载荷的相关性(> 0.6),PQ和CS在恢复峰强度和再现OPLS-DA模型的真实分类特征方面表现最好,而与光谱噪声水平无关。我们的研究结果表明,性能在很大程度上取决于数据集中的噪声水平,而稀释因子的影响可以忽略不计。对于有效的NMR代谢组学数据集,还确定了20%的最小允许噪声水平。
Failure to properly account for normal systematic variations in OMICS datasets may result in misleading biological conclusions. Accordingly, normalization is a necessary step in the proper preprocessing of OMICS datasets. In this regards, an optimal normalization method will effectively reduce unwanted biases and increase the accuracy of downstream quantitative analyses. But, it is currently unclear which normalization method is best since each algorithm addresses systematic noise in different ways.Determine an optimal choice of a normalization method for the preprocessing of metabolomics datasets.Nine MVAPACK normalization algorithms were compared with simulated and experimental NMR spectra modified with added Gaussian noise and random dilution factors. Methods were evaluated based on an ability to recover the intensities of the true spectral peaks and the reproducibility of true classifying features from orthogonal projections to latent structures-discriminant analysis model (OPLS-DA).Most normalization methods (except histogram matching) performed equally well at modest levels of signal variance. Only probabilistic quotient (PQ) and constant sum (CS) maintained the highest level of peak recovery (> 67%) and correlation with true loadings (> 0.6) at maximal noise.PQ and CS performed the best at recovering peak intensities and reproducing the true classifying features for an OPLS-DA model regardless of spectral noise level. Our findings suggest that performance is largely determined by the level of noise in the dataset, while the effect of dilution factors was negligible. A minimal allowable noise level of 20% was also identified for a valid NMR metabolomics dataset.