Comparing different supervised machine learning algorithms for disease prediction

Comparing different supervised machine learning algorithms for disease prediction
复制标题

DOI:
10.1186/s12911-019-1004-8
复制
发表时间:
2019-12-21
影响因子:
3.5
通讯作者:
Moni, Mohammad Ali
Moni, Mohammad Ali
中科院分区:
医学3区
文献类型:
--
作者:
Uddin, Shahadat;Khan, Arif;Moni, Mohammad Ali

文献摘要

被引文献

相似文献

背景有监督机器学习算法已经成为数据挖掘领域的主流方法。使用健康数据的疾病预测最近显示了这些方法的潜在应用领域。本研究旨在确定不同类型的监督机器学习算法之间的关键趋势,以及它们在疾病风险预测中的性能和用途。方法在这项研究中,进行了广泛的研究工作,以确定那些在单一疾病预测中应用了多种监督机器学习算法的研究。两个数据库(即,Scopus和PubMed)检索不同类型的检索项。因此,我们总共选择了48篇文章来比较用于疾病预测的变体监督机器学习算法。结果支持向量机(SVM)算法的应用频率最高(29项研究),其次是朴素贝叶斯算法(23项研究)。然而,随机森林(RF)算法表现出上级的准确性比较。在应用射频的17项研究中,其中9项显示出最高的准确性,即,百分之五十三其次是SVM,在41%的研究中名列前茅。结论这项研究提供了一个广泛的概述不同变体的监督机器学习算法的疾病预测的相对性能。相对性能的这一重要信息可用于帮助研究人员为他们的研究选择适当的监督机器学习算法。
Background Supervised machine learning algorithms have been a dominant method in the data mining field. Disease prediction using health data has recently shown a potential application area for these methods. This study ai7ms to identify the key trends among different types of supervised machine learning algorithms, and their performance and usage for disease risk prediction. Methods In this study, extensive research efforts were made to identify those studies that applied more than one supervised machine learning algorithm on single disease prediction. Two databases (i.e., Scopus and PubMed) were searched for different types of search items. Thus, we selected 48 articles in total for the comparison among variants supervised machine learning algorithms for disease prediction. Results We found that the Support Vector Machine (SVM) algorithm is applied most frequently (in 29 studies) followed by the Naive Bayes algorithm (in 23 studies). However, the Random Forest (RF) algorithm showed superior accuracy comparatively. Of the 17 studies where it was applied, RF showed the highest accuracy in 9 of them, i.e., 53%. This was followed by SVM which topped in 41% of the studies it was considered. Conclusion This study provides a wide overview of the relative performance of different variants of supervised machine learning algorithms for disease prediction. This important information of relative performance can be used to aid researchers in the selection of an appropriate supervised machine learning algorithm for their studies.