Machine Learning and Data Mining Methods in Diabetes Research.

Machine Learning and Data Mining Methods in Diabetes Research.
复制标题

糖尿病研究中的机器学习和数据挖掘方法。

DOI:
10.1016/j.csbj.2016.12.005
复制
发表时间:
2017
影响因子:
6
通讯作者:
Chouvarda I
Chouvarda I
中科院分区:
生物学2区
文献类型:
--
作者:
Kavakiotis I;Tsave O;Salifoglou A;Maglaveras N;Vlahavas I;Chouvarda I

文献摘要

被引文献

相似文献

生物技术和健康科学的显著进步导致了大量数据的产生,例如从大型电子健康记录(EHR)产生的高通量基因数据和临床信息。为此,机器学习和数据挖掘方法在生物科学中的应用目前比以往任何时候都更加重要和不可或缺,以努力将所有可用信息智能地转化为有价值的知识。糖尿病(DM)是指在全球范围内对人类健康造成巨大压力的一组代谢紊乱。广泛研究糖尿病的各个方面(诊断、病理生理学、治疗等)导致了海量数据的产生。本研究的目的是对机器学习、数据挖掘技术和工具在糖尿病研究领域中的应用进行系统的回顾,涉及a)预测和诊断,b)糖尿病并发症,c)遗传背景和环境,以及e)健康护理和管理,其中第一类似乎是最受欢迎的。使用了广泛的机器学习算法。总的来说,85%的被使用的学习方法是有监督的学习方法,15%的人是无监督的学习方法,更具体地说,是关联规则。支持向量机作为最成功、应用最广泛的算法应运而生。在数据类型方面,主要使用临床数据集。在选定的文章中的标题应用计划提取有价值的知识的有用性,导致新的假设,目标是更深入的理解和进一步的研究DM。
The remarkable advances in biotechnology and health sciences have led to a significant production of data, such as high throughput genetic data and clinical information, generated from large Electronic Health Records (EHRs). To this end, application of machine learning and data mining methods in biosciences is presently, more than ever before, vital and indispensable in efforts to transform intelligently all available information into valuable knowledge. Diabetes mellitus (DM) is defined as a group of metabolic disorders exerting significant pressure on human health worldwide. Extensive research in all aspects of diabetes (diagnosis, etiopathophysiology, therapy, etc.) has led to the generation of huge amounts of data. The aim of the present study is to conduct a systematic review of the applications of machine learning, data mining techniques and tools in the field of diabetes research with respect to a) Prediction and Diagnosis, b) Diabetic Complications, c) Genetic Background and Environment, and e) Health Care and Management with the first category appearing to be the most popular. A wide range of machine learning algorithms were employed. In general, 85% of those used were characterized by supervised learning approaches and 15% by unsupervised ones, and more specifically, association rules. Support vector machines (SVM) arise as the most successful and widely used algorithm. Concerning the type of data, clinical datasets were mainly used. The title applications in the selected articles project the usefulness of extracting valuable knowledge leading to new hypotheses targeting deeper understanding and further investigation in DM.