Quantitative Comparison of Statistical Methods for Analyzing Human Metabolomics Data.

Quantitative Comparison of Statistical Methods for Analyzing Human Metabolomics Data.
复制标题

用于分析人类代谢组学数据的统计方法的定量比较。

DOI:
10.3390/metabo12060519
复制
发表时间:
2022-06-04
期刊:
影响因子:
4.1
通讯作者:
--
中科院分区:
生物学3区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

新兴技术现在允许在越来越多的生物样品中对数千种小分子代谢物进行基于质谱的分析(“代谢组学”)。虽然为深入了解人类疾病的发病机制提供了很大的希望,但尚未建立标准方法来统计分析与临床表型(包括疾病结局)相关的日益复杂的高维人类代谢组学数据。为了确定最佳的分析方法,我们正式比较了一系列代谢组学数据集类型的传统和更新的统计学习方法。在来自大型人群的模拟和实验代谢组学数据中,我们观察到随着研究受试者数量的增加,单变量与多变量方法相比导致明显更高的错误发现率,如与结果直接相关的代谢物和与结果无关的代谢物之间的显著相关性所示。虽然从严格的统计学意义上说,这种关联的频率较高不会被认为是错误的,但它可能被认为是生物学上的信息较少。在测定的代谢物数量增加的情况下,如在非靶向代谢组学与靶向代谢组学的测量中,多变量方法在一系列统计操作特征中表现得特别有利。在包括数千种代谢物测量的非靶向代谢组学数据集中,稀疏多变量模型表现出更高的选择性和更低的虚假关系的可能性。当代谢物的数量类似于或超过研究受试者的数量时,这在相对较小的队列的非靶向代谢组学分析中很常见,稀疏多变量模型表现出最稳健的统计功效,结果更一致。这些发现对人类疾病的代谢组学分析具有重要意义。
Emerging technologies now allow for mass spectrometry-based profiling of thousands of small molecule metabolites (‘metabolomics’) in an increasing number of biosamples. While offering great promise for insight into the pathogenesis of human disease, standard approaches have not yet been established for statistically analyzing increasingly complex, high-dimensional human metabolomics data in relation to clinical phenotypes, including disease outcomes. To determine optimal approaches for analysis, we formally compare traditional and newer statistical learning methods across a range of metabolomics dataset types. In simulated and experimental metabolomics data derived from large population-based human cohorts, we observe that with an increasing number of study subjects, univariate compared to multivariate methods result in an apparently higher false discovery rate as represented by substantial correlation between metabolites directly associated with the outcome and metabolites not associated with the outcome. Although the higher frequency of such associations would not be considered false in the strict statistical sense, it may be considered biologically less informative. In scenarios wherein the number of assayed metabolites increases, as in measures of nontargeted versus targeted metabolomics, multivariate methods performed especially favorably across a range of statistical operating characteristics. In nontargeted metabolomics datasets that included thousands of metabolite measures, sparse multivariate models demonstrated greater selectivity and lower potential for spurious relationships. When the number of metabolites was similar to or exceeded the number of study subjects, as is common with nontargeted metabolomics analysis of relatively small cohorts, sparse multivariate models exhibited the most-robust statistical power with more consistent results. These findings have important implications for metabolomics analysis in human disease.
DOI: 10.1021/acs.analchem.7b02380
发表时间: 2017-10-03
影响因子: 7.4
作者:
Mahieu NG;Patti GJ
通讯作者: Patti GJ
DOI: 10.1002/cem.785
发表时间: 2003-03-01
影响因子: 2.4
作者:
Barker, M;Rayens, W
通讯作者: Rayens, W
DOI: 10.1002/jms.3780
发表时间: 2016-08
期刊: Journal of mass spectrometry : JMS
影响因子: --
作者:
Barnes S;Benton HP;Casazza K;Cooper SJ;Cui X;Du X;Engler J;Kabarowski JH;Li S;Pathmasiri W;Prasain JK;Renfrow MB;Tiwari HK
通讯作者: Tiwari HK
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.3389/fbioe.2015.00023
发表时间: 2015
影响因子: 5.7
作者:
Alonso A;Marsal S;Julià A
通讯作者: Julià A