The development and validation of a non-invasive prediction model of hyperuricemia based on modifiable risk factors: baseline findings of a health examination population cohort

The development and validation of a non-invasive prediction model of hyperuricemia based on modifiable risk factors: baseline findings of a health examination population cohort
复制标题

DOI:
10.1039/d3fo01363d
复制
发表时间:
2023-05-17
期刊:
影响因子:
6.1
通讯作者:
He,Huijing
He,Huijing
中科院分区:
农林科学1区
文献类型:
--
作者:
Chen,Shuo;Han,Wei;He,Huijing

文献摘要

被引文献

相似文献

本研究的目的是建立一个简单的,非侵入性的高尿酸血症的风险预测模型,在中国成年人的基础上修改的危险因素。2020-2021年,在北京市健康体检人群中开展了北京市健康管理队列(BHMC)基线调查。收集了各种生活方式风险因素,包括饮食模式和习惯,吸烟,饮酒,睡眠时间和使用手机。我们使用三种机器学习技术开发了高尿酸血症预测模型,即逻辑回归(LR),随机森林(RF)和XGBoost。比较了这三种方法在区分、校准和临床适用性方面的性能。决策曲线分析(DCA)用于评估模型的临床实用性。共有74 050人被纳入研究,其中55 537人(75%)被随机选择到训练集中,其他18 513人(25%)被纳入验证集中。   HUA患病率男性为38.43%,女性为13.29%。XGBoost模型比LR和RF模型具有更好的性能。LR、RF和XGBoost模型训练集中的曲线下面积(AUC)(95% CI)分别为0.754(0.750-0.757)、0.844(0.841-0.846)和0.854(0.851-0.856)。XGBoost模型的分类准确率为0.774,高于logistic模型(0.592)和RF模型(0.767)。LR、RF和XGBoost模型验证集中的AUC(95% CI)值分别为0.758(0.749-0.765)、0.809(0.802-0.816)和0.820(0.813-0.827)。如DCA曲线所示,所有三种模式都可以在适当的阈值概率内带来净效益。XGBoost具有更好的区分度和准确性。模型中包含的各种可改变的危险因素有助于促进HUA高危人群的容易识别和生活方式干预。
This study aims to establish a simple and non-invasive risk prediction model for hyperuricemia in Chinese adults based on modifiable risk factors. In 2020–2021, the baseline survey of the Beijing Health Management Cohort (BHMC) was conducted in Beijing city among the health examination population. Diverse life-style risk factors including dietary patterns and habits, cigarette smoking, alcohol intake, sleep duration and cell-phone use were collected. We developed hyperuricemia prediction models using three machine-learning techniques, namely logistic regression (LR), random forest (RF), and XGBoost. Performances in discrimination, calibration, and clinical applicability of the three methods were compared. Decision curve analysis (DCA) was used to assess the model's clinical usefulness. A total of 74 050 people were included in the study, of whom 55 537 (75%) were randomly selected into the training set and the other 18 513 (25%) were in the validation set. The prevalence of HUA was 38.43% in men and 13.29% in women. The XGBoost model has better performance than the LR and RF models. The area under the curve (AUC) (95% CI) in the training set for the LR, RF and XGBoost models were 0.754 (0.750–0.757), 0.844 (0.841–0.846) and 0.854 (0.851–0.856), respectively. The XGBoost model had a higher classification accuracy of 0.774 than the logistic (0.592) and RF (0.767) models. The AUC (95% CI) values in the validation set for the LR, RF and XGBoost models were 0.758 (0.749–0.765), 0.809 (0.802–0.816) and 0.820 (0.813–0.827), respectively. As demonstrated by the DCA curves, all the three models could bring net benefits within the appropriate threshold probability. XGBoost had better discrimination and accuracy. Various modifiable risk factors included in the model were helpful in facilitating the easy identification and life-style interventions of the HUA high-risk population.