425Artificial Intelligence Approaches to Type 2 Diabetes Risk Prediction and Exploration of Predictive Factors

425Artificial Intelligence Approaches to Type 2 Diabetes Risk Prediction and Exploration of Predictive Factors
复制标题

4252型糖尿病风险预测的人工智能方法和预测因素的探索

DOI:
10.1093/ije/dyab168.515
复制
发表时间:
2021
影响因子:
7.7
通讯作者:
Yamagata Zentaro
Yamagata Zentaro
中科院分区:
医学1区
文献类型:
--
作者:
Ooka Tadao;Yokomichi Hiroshi;Yamagata Zentaro

文献摘要

相似文献

将人工智能纳入流行病学,特别是在数据解释方面存在主要障碍。因此,我们研究了高度可解释的机器学习方法——随机森林(RF)和稀疏逻辑回归(SLR)——在大规模健康检查数据集上的应用,并研究了使用这些方法创建预测模型的优势。方法本研究涉及1999年至2018年在日本接受健康检查的392791名参与者。接受糖尿病治疗或HbA1c水平为6.5%或更高的参与者被排除在外。研究的客观变量是5年内2型糖尿病的发病情况。每个预测模型是使用连续三年的26个健康状态项目创建的。我们检验了三种分析方法来比较它们的预测能力:RF、SLR和作为常规方法的多元逐步逻辑回归(MSLR)。在RF分析中计算变量重要性(VI),在SLR和MSLR分析中计算标准回归系数(SRC)。结果SLR模型预测准确率最高(AUC:0.955), RF模型次之(AUC:0.949), MSLR模型次之(AUC:0.939)。RF模型测量血糖、糖化血红蛋白、身高、红细胞和天冬氨酸转氨酶,具有较高的预测能力。在SLR模型中,HbA1c、血糖、收缩压、hdl -胆固醇和年龄均有较高的SRC。结论机器学习技术能够比现有方法更准确地预测糖尿病风险,并提出了识别相关预测因子的新方法。将机器学习方法应用于健康检查数据,在保持数据可解释性的同时,在预测2型糖尿病方面达到了很高的准确性。
BackgroundMajor barriers exist in incorporating artificial intelligence into epidemiology, particularly in data interpretation. Thus, we examined the application of highly interpretable machine-learning methods— Random Forest (RF) and Sparse Logistic Regression (SLR)— to a large-scale health check-up dataset, examining the advantages of creating prediction models using these.MethodsThis study involved 392,791 participants who underwent healthcare checkups in Japan from 1999 to 2018. Participants who received diabetes treatment, or had an HbA1c level of 6.5% or higher, were excluded. The objective variable examined was type 2 diabetes onset over five years. Each prediction model was created using 26 health status items over three consecutive years. We examined three analytical methods to compare their predictive powers: RF, SLR, and a multivariate stepwise logistic regression (MSLR) as a conventional method. Variable Importance (VI) was calculated in the RF analysis, with Standard Regression Coefficients (SRC) being calculated in the SLR and MSLR analyses.ResultsPredictive accuracy is highest in the SLR model (AUC:0.955), followed by the RF model (AUC:0.949), and then the MSLR model (AUC:0.939). The RF model measures blood glucose, HbA1c, height, red blood cells, and aspartate transaminase with a higher predictive power. In the SLR model, HbA1c, blood glucose, systolic blood pressure, HDL-Cholesterol, and age have higher SRC.ConclusionsMachine learning techniques enable more accurate diabetes risk predictions than existing methods and suggest new ways of identifying associated predictors.Key messagesApplying machine-learning methods to health check-up data achieves a high accuracy in predicting type 2 diabetes while maintaining data interpretability.