A comparative evaluation of the generalised predictive ability of eight machine learning algorithms across ten clinical metabolomics data sets for binary classification

A comparative evaluation of the generalised predictive ability of eight machine learning algorithms across ten clinical metabolomics data sets for binary classification
复制标题

DOI:
10.1007/s11306-019-1612-4
复制
发表时间:
2019-12-01
期刊:
影响因子:
3.6
通讯作者:
Broadhurst, David, I
Broadhurst, David, I
中科院分区:
医学3区
文献类型:
--
作者:
Mendez, Kevin M.;Reinke, Stacey N.;Broadhurst, David, I

文献摘要

被引文献

相似文献

前言代谢组学正越来越多地被用于疾病诊断、预后和风险预测的临床环境中。机器学习算法在构建多变量代谢物预测中尤为重要。从历史上看,偏最小二乘回归一直是二进制分类的黄金标准。随机森林(RF)、核支持向量机(SVM)和人工神经网络(ANN)等非线性机器学习方法可能更适合对可能的非线性代谢物协方差进行建模,从而提供更好的预测模型。目的我们假设,对于利用代谢组学数据进行二分类的方法,与线性方法相比,尤其是与当前的金标准偏最小二乘判别分析相比,非线性机器学习方法将提供更好的泛化预测能力。方法我们在10个公开可用的临床代谢组学数据集上比较了8种原型机器学习算法的总体预测性能。这些算法是用Python语言实现的。所有代码和结果都已作为Jupyter笔记本公开提供。结果在所有的数据集上,支持向量机和神经网络相对于偏最小二乘法的预测能力只有轻微的改善。射频性能相对较差。使用现成的Bootstrap可信区间提供了一种模型预测的不确定性度量,因此观察到代谢组学数据的质量对普遍表现的影响比模型选择更大。结论与机器学习算法的选择相比,数据集的大小和性能指标的选择对广义预测性能的影响更大。
Introduction Metabolomics is increasingly being used in the clinical setting for disease diagnosis, prognosis and risk prediction. Machine learning algorithms are particularly important in the construction of multivariate metabolite prediction. Historically, partial least squares (PLS) regression has been the gold standard for binary classification. Nonlinear machine learning methods such as random forests (RF), kernel support vector machines (SVM) and artificial neural networks (ANN) may be more suited to modelling possible nonlinear metabolite covariance, and thus provide better predictive models. Objectives We hypothesise that for binary classification using metabolomics data, non-linear machine learning methods will provide superior generalised predictive ability when compared to linear alternatives, in particular when compared with the current gold standard PLS discriminant analysis. Methods We compared the general predictive performance of eight archetypal machine learning algorithms across ten publicly available clinical metabolomics data sets. The algorithms were implemented in the Python programming language. All code and results have been made publicly available as Jupyter notebooks. Results There was only marginal improvement in predictive ability for SVM and ANN over PLS across all data sets. RF performance was comparatively poor. The use of out-of-bag bootstrap confidence intervals provided a measure of uncertainty of model prediction such that the quality of metabolomics data was observed to be a bigger influence on generalised performance than model choice. Conclusion The size of the data set, and choice of performance metric, had a greater influence on generalised predictive performance than the choice of machine learning algorithm.