Automated machine learning and explainable AI (AutoML-XAI) for metabolomics: improving cancer diagnostics.

Automated machine learning and explainable AI (AutoML-XAI) for metabolomics: improving cancer diagnostics.
复制标题

用于代谢组学的自动化机器学习和可解释的人工智能 (AutoML-XAI):改善癌症诊断。

DOI:
10.1101/2023.10.26.564244
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Fernández,FacundoM
Fernández,FacundoM
中科院分区:
--
文献类型:
--
作者:
Bifarin,OlatomiwaO;Fernández,FacundoM

文献摘要

相似文献

代谢组学产生复杂的数据,需要先进的计算方法来产生生物学洞察力。虽然机器学习 (ML) 前景广阔,但选择最佳算法和调整超参数的挑战仍然存在,尤其是对于非专家而言。自动机器学习(AutoML)可以简化这个过程;然而,可解释性问题可能仍然存在。这项研究引入了一个统一的流程,将 AutoML 与可解释的人工智能 (XAI) 技术相结合,以优化代谢组学分析。我们在两个数据集上测试了我们的方法:肾细胞癌(RCC)尿液代谢组学和卵巢癌(OC)血清代谢组学。使用 Auto-sklearn 的 AutoML 在区分 RCC 和健康对照以及 OC 患者和其他妇科癌症患者方面超越了 SVM 和 k-Nearest Neighbors 等独立 ML 算法。从未见过的测试集获得的 RCC 的 AUC 分数为 0.97,OC 的 AUC 分数为 0.85,凸显了 Auto-sklearn 的有效性。重要的是,在大多数考虑的指标上,Auto-sklearn 利用算法和集成技术的组合展示了更好的分类性能。 Shapley Additive Explanations (SHAP) 提供了特征重要性的全球排名,将二丁胺和神经节苷脂 GM(d34:1) 分别确定为 RCC 和 OC 的首要区分代谢物。瀑布图通过说明每种代谢物对个体预测的影响提供了局部解释。依赖性图突出了代谢物相互作用,例如 RCC 中马尿酸及其衍生物之一之间的联系,以及 OC 中 GM3(d34:1) 和 GM3(18:1_16:0) 之间的联系,暗示了潜在的机制关系。通过决策图,进行了详细的误差分析,对比了正确分类样本和错误分类样本的特征重要性。从本质上讲,我们的流程强调协调 AutoML 和 XAI 的重要性,促进简化的 ML 应用和提高代谢组学数据科学的可解释性。
Metabolomics generates complex data necessitating advanced computational methods for generating biological insight. While machine learning (ML) is promising, the challenges of selecting the best algorithms and tuning hyperparameters, particularly for nonexperts, remain. Automated machine learning (AutoML) can streamline this process; however, the issue of interpretability could persist. This research introduces a unified pipeline that combines AutoML with explainable AI (XAI) techniques to optimize metabolomics analysis. We tested our approach on two data sets: renal cell carcinoma (RCC) urine metabolomics and ovarian cancer (OC) serum metabolomics. AutoML, using Auto-sklearn, surpassed standalone ML algorithms like SVM and k-Nearest Neighbors in differentiating between RCC and healthy controls, as well as OC patients and those with other gynecological cancers. The effectiveness of Auto-sklearn is highlighted by its AUC scores of 0.97 for RCC and 0.85 for OC, obtained from the unseen test sets. Importantly, on most of the metrics considered, Auto-sklearn demonstrated a better classification performance, leveraging a mix of algorithms and ensemble techniques. Shapley Additive Explanations (SHAP) provided a global ranking of feature importance, identifying dibutylamine and ganglioside GM(d34:1) as the top discriminative metabolites for RCC and OC, respectively. Waterfall plots offered local explanations by illustrating the influence of each metabolite on individual predictions. Dependence plots spotlighted metabolite interactions, such as the connection between hippuric acid and one of its derivatives in RCC, and between GM3(d34:1) and GM3(18:1_16:0) in OC, hinting at potential mechanistic relationships. Through decision plots, a detailed error analysis was conducted, contrasting feature importance for correctly versus incorrectly classified samples. In essence, our pipeline emphasizes the importance of harmonizing AutoML and XAI, facilitating both simplified ML application and improved interpretability in metabolomics data science.