Statistical considerations and database limitations in NMR-based metabolic profiling studies

Statistical considerations and database limitations in NMR-based metabolic profiling studies
复制标题

DOI:
10.1007/s11306-023-02027-5
复制
发表时间:
2023-06
期刊:
影响因子:
3.6
通讯作者:
Imani L. Ross;Julie A. Beardslee;Maria M. Steil;Tafadzwa Chihanga;M. Kennedy
Imani L. Ross;Julie A. Beardslee;Maria M. Steil;Tafadzwa Chihanga;M. Kennedy
中科院分区:
医学3区
文献类型:
--
作者:
Imani L. Ross;Julie A. Beardslee;Maria M. Steil;Tafadzwa Chihanga;M. Kennedy

文献摘要

相似文献

简介基于 NMR 的代谢谱研究的解释和分析受到实质上不完整的商业和学术数据库的限制。统计显着性检验(包括 p 值、VIP 分数、AUC 值和 FC 值)可能在很大程度上不一致。统计分析之前的数据标准化可能会导致错误的结果。目标目标是 (1) 定量评估基于 NMR 的代表性代谢分析数据集中的 p 值、VIP 分数、AUC 值和 FC 值之间的一致性,(2) 评估数据标准化如何影响统计显着性结果,(3) 使用常用数据库确定共振峰分配完成潜力,以及 (4) 分析这些数据库中代谢物空间的交叉和独特性。方法 P 值、VIP 分数、AUC值和 FC 值及其对数据标准化的依赖性是在胰腺癌原位小鼠模型和两种人胰腺癌细胞系中确定的。使用 Chenomx、人类代谢物数据库 (HMDB) 和 COLMAR 数据库评估共振分配的完整性。量化了数据库的交叉性和唯一性。结果与 VIP 或 FC 值相比,P 值和 AUC 值密切相关。具有统计显着性的箱的分布在很大程度上取决于数据集是否标准化。 40-45% 的峰没有数据库匹配或数据库匹配不明确。每个数据库中有 9-22% 的代谢物是独特的。结论代谢组学数据统计分析缺乏一致性可能会导致误导性或不一致的解释。数据标准化会对统计分析产生很大影响,因此应该合理化。大约 40% 的峰值分配在当前数据库中仍然不明确或不可能。一维和二维数据库应保持一致,以最大限度地提高代谢物分配的置信度和验证。
IntroductionInterpretation and analysis of NMR-based metabolic profiling studies is limited by substantially incomplete commercial and academic databases. Statistical significance tests, including p-values, VIP scores, AUC values and FC values, can be largely inconsistent. Data normalization prior to statistical analysis can cause erroneous outcomes.ObjectivesThe objectives were (1) to quantitatively assess consistency among p-values, VIP scores, AUC values and FC values in representative NMR-based metabolic profiling datasets, (2) to assess how data normalization can impact statistical significance outcomes, (3) to determine resonance peak assignment completion potential using commonly used databases and (4) to analyze intersection and uniqueness of metabolite space in these databases.MethodsP-values, VIP scores, AUC values and FC values, and their dependence on data normalization, were determined in orthotopic mouse model of pancreatic cancer and two human pancreatic cancer cell lines. Completeness of resonance assignments were evaluated using Chenomx, the human metabolite database (HMDB) and the COLMAR database. The intersection and uniqueness of the databases was quantified.ResultsP-values and AUC values were strongly correlated compared to VIP or FC values. Distributions of statistically significant bins depended strongly on whether or not datasets were normalized. 40–45% of peaks had either no or ambiguous database matches. 9–22% of metabolites were unique to each database.ConclusionsLack of consistency in statistical analyses of metabolomics data can lead to misleading or inconsistent interpretation. Data normalization can have large effects on statistical analysis and should be justified. About 40% of peak assignments remain ambiguous or impossible with current databases. 1D and 2D databases should be made consistent to maximize metabolite assignment confidence and validation.