425Artificial Intelligence Approaches to Type 2 Diabetes Risk Prediction and Exploration of Predictive Factors
425Artificial Intelligence Approaches to Type 2 Diabetes Risk Prediction and Exploration of Predictive Factors
复制标题
4252型糖尿病风险预测的人工智能方法和预测因素的探索
DOI:
10.1093/ije/dyab168.515
复制
发表时间:
2021
影响因子:
7.7
通讯作者:
Yamagata Zentaro
中科院分区:
文献类型:
--
作者:
Ooka Tadao;Yokomichi Hiroshi;Yamagata Zentaro
BackgroundMajor barriers exist in incorporating artificial intelligence into epidemiology, particularly in data interpretation. Thus, we examined the application of highly interpretable machine-learning methods— Random Forest (RF) and Sparse Logistic Regression (SLR)— to a large-scale health check-up dataset, examining the advantages of creating prediction models using these.MethodsThis study involved 392,791 participants who underwent healthcare checkups in Japan from 1999 to 2018. Participants who received diabetes treatment, or had an HbA1c level of 6.5% or higher, were excluded. The objective variable examined was type 2 diabetes onset over five years. Each prediction model was created using 26 health status items over three consecutive years. We examined three analytical methods to compare their predictive powers: RF, SLR, and a multivariate stepwise logistic regression (MSLR) as a conventional method. Variable Importance (VI) was calculated in the RF analysis, with Standard Regression Coefficients (SRC) being calculated in the SLR and MSLR analyses.ResultsPredictive accuracy is highest in the SLR model (AUC:0.955), followed by the RF model (AUC:0.949), and then the MSLR model (AUC:0.939). The RF model measures blood glucose, HbA1c, height, red blood cells, and aspartate transaminase with a higher predictive power. In the SLR model, HbA1c, blood glucose, systolic blood pressure, HDL-Cholesterol, and age have higher SRC.ConclusionsMachine learning techniques enable more accurate diabetes risk predictions than existing methods and suggest new ways of identifying associated predictors.Key messagesApplying machine-learning methods to health check-up data achieves a high accuracy in predicting type 2 diabetes while maintaining data interpretability.