Pathway analysis in metabolomics: Recommendations for the use of over-representation analysis.

Pathway analysis in metabolomics: Recommendations for the use of over-representation analysis.
复制标题

DOI:
10.1371/journal.pcbi.1009105
复制
发表时间:
2021-09
影响因子:
4.3
通讯作者:
Ebbels T
Ebbels T
中科院分区:
生物学2区
文献类型:
--
作者:
Wieder C;Frainay C;Poupin N;Rodríguez-Mier P;Vinson F;Cooke J;Lai RP;Bundy JG;Jourdan F;Ebbels T

文献摘要

参考文献

被引文献

相似文献

过代表性分析(ORA)是代谢组学数据集功能解释中最常用的途径分析方法之一。尽管ORA在代谢组学中广泛使用,但社区缺乏详细说明其最佳实践使用的指南。许多因素对结果有明显的影响,但迄今为止,它们的影响很少得到系统的关注。使用五个公开可用的数据集,我们证明了参数的变化,如背景集,差异代谢物选择方法和所用的途径数据库,可以导致截然不同的ORA结果。例如,使用非测定特异性背景组会导致大量假阳性途径。使用三种最流行的代谢途径数据库(KEGG,Reactome和BioCyc)进行评估的途径数据库选择导致显著富集途径的数量和功能存在巨大差异。特定于代谢组学数据的因素,如化合物鉴定的可靠性和不同分析平台的化学偏倚也会影响ORA结果。模拟的代谢物错误识别率低至4%,导致所有数据集的假阳性途径增加和真正重要途径的损失。我们的研究结果对ORA用户以及使用替代途径分析方法的用户有一些实际意义。我们提供了一套在代谢组学中使用ORA的建议,以及一套最低限度的报告指南,作为代谢组学途径分析标准化的第一步。代谢组学是一个快速发展的研究领域,涉及生物体内小分子的分析。它使研究人员能够了解生物状态(如健康或疾病)对细胞生物化学的影响,并具有广泛的应用,从生物标志物发现和医疗保健中的个性化药物到农业中的作物保护和粮食安全。途径分析有助于了解哪些生物途径,代表执行特定功能的分子集合,可能参与对疾病表型或药物治疗的反应。过度代表性分析(ORA)可能是代谢组学领域最常用的途径分析方法。然而,ORA可以根据所使用的输入数据和参数给出截然不同的结果。在这里,我们已经建立了这些因素对ORA结果的影响,使用应用于五个真实世界数据集的计算修改。根据我们的研究结果,我们为研究界提供了一套最佳实践建议,不仅适用于ORA,也适用于其他途径分析方法,以帮助确保结果的可靠性和重现性。
Over-representation analysis (ORA) is one of the commonest pathway analysis approaches used for the functional interpretation of metabolomics datasets. Despite the widespread use of ORA in metabolomics, the community lacks guidelines detailing its best-practice use. Many factors have a pronounced impact on the results, but to date their effects have received little systematic attention. Using five publicly available datasets, we demonstrated that changes in parameters such as the background set, differential metabolite selection methods, and pathway database used can result in profoundly different ORA results. The use of a non-assay-specific background set, for example, resulted in large numbers of false-positive pathways. Pathway database choice, evaluated using three of the most popular metabolic pathway databases (KEGG, Reactome, and BioCyc), led to vastly different results in both the number and function of significantly enriched pathways. Factors that are specific to metabolomics data, such as the reliability of compound identification and the chemical bias of different analytical platforms also impacted ORA results. Simulated metabolite misidentification rates as low as 4% resulted in both gain of false-positive pathways and loss of truly significant pathways across all datasets. Our results have several practical implications for ORA users, as well as those using alternative pathway analysis methods. We offer a set of recommendations for the use of ORA in metabolomics, alongside a set of minimal reporting guidelines, as a first step towards the standardisation of pathway analysis in metabolomics. Metabolomics is a rapidly growing field of study involving the profiling of small molecules within an organism. It allows researchers to understand the effects of biological status (such as health or disease) on cellular biochemistry, and has wide-ranging applications, from biomarker discovery and personalised medicine in healthcare to crop protection and food security in agriculture. Pathway analysis helps to understand which biological pathways, representing collections of molecules performing a particular function, may be involved in response to a disease phenotype, or drug treatment, for example. Over-representation analysis (ORA) is perhaps the most common pathway analysis method used in the metabolomics community. However, ORA can give drastically different results depending on the input data and parameters used. Here, we have established the effects of these factors on ORA results using computational modifications applied to five real-world datasets. Based on our results, we offer the research community a set of best-practice recommendations applicable not only to ORA but also to other pathway analysis methods to help ensure the reliability and reproducibility of results.
代谢组学和途径分析表征妊娠奶牛在 AI 后第 17 天和第 45 天的代谢变化
DOI: 10.1038/s41598-018-23983-2
发表时间: 2018-04-13
期刊: Scientific reports
影响因子: 4.6
作者:
Guo YS;Tao JZ
通讯作者: Tao JZ
DOI: 10.1146/annurev-biochem-061516-044952
发表时间: 2017-06-20
影响因子: 16.6
作者:
Lu W;Su X;Klein MS;Lewis IA;Fiehn O;Rabinowitz JD
通讯作者: Rabinowitz JD
DOI: 10.1038/s41540-018-0078
发表时间: 2018-12-13
影响因子: 4
作者:
Domingo-Fernandez, Daniel;Hoyt, Charles Tapley;Hofmann-Apitius, Martin
通讯作者: Hofmann-Apitius, Martin
DOI: 10.1093/bioinformatics/btt547
发表时间: 2013-12-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cokelaer T;Pultz D;Harder LM;Serra-Musach J;Saez-Rodriguez J
通讯作者: Saez-Rodriguez J
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y