Comparison of machine-learning and logistic regression models to predict 30-day unplanned readmission: a development and validation study
Comparison of machine-learning and logistic regression models to predict 30-day unplanned readmission: a development and validation study
复制标题
机器学习和逻辑回归模型预测 30 天计划外再入院的比较:一项开发和验证研究
DOI:
10.1101/2023.05.06.23289569
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Nanako Tamiya
中科院分区:
文献类型:
--
作者:
Masao Iwagami;Ryota Inokuchi;Eiryo Kawakami;Tomohide Yamada;Atsushi Goto;Toshiki Kuno;Yohei Hashimoto;Nobuaki Michihata;Tadahiro Goto;Tomohiro Shinozaki;Yu Sun;Yuta Taniguchi;Jun Komiyama;Kazuaki Uda;Toshikazu Abe;Nanako Tamiya
It is expected but unknown whether machine-learning models can outperform regression models, such as a logistic regression (LR) model, especially when the number and types of predictor variables increase in electronic health records (EHRs). We aimed to compare the predictive performance of gradient-boosted decision tree (GBDT), random forest (RF), deep neural network (DNN), and LR with the least absolute shrinkage and selection operator (LR-LASSO) for unplanned readmission. We used EHRs of patients discharged alive from 38 hospitals in 2015–2017 for derivation and in 2018 for validation, including basic characteristics, diagnosis, surgery, procedure, and drug codes, and blood-test results. The outcome was 30-day unplanned readmission. We created six patterns of data tables having different numbers of binary variables (that ≥5% or ≥1% of patients or ≥10 patients had) with and without blood-test results. For each pattern of data tables, we used the derivation data to establish the machine-learning and LR models, and used the validation data to evaluate the performance of each model. The incidence of outcome was 6.8% (23,108/339,513 discharges) and 6.4% (7,507/118,074 discharges) in the derivation and validation datasets, respectively. For the first data table with the smallest number of variables (102 variables that ≥5% of patients had, without blood-test results), the c-statistic was highest for GBDT (0.740), followed by RF (0.734), LR-LASSO (0.720), and DNN (0.664). For the last data table with the largest number of variables (1543 variables that ≥10 patients had, including blood-test results), the c-statistic was highest for GBDT (0.764), followed by LR-LASSO (0.755), RF (0.751), and DNN (0.720), suggesting that the difference between GBDT and LR-LASSO was small and their 95% confidence intervals overlapped. In conclusion, GBDT generally outperformed LR-LASSO to predict unplanned readmission, but the difference of c-statistic became smaller as the number of variables was increased and blood-test results were used.
登录
查看更多内容
DOI:
10.1542/9781610021074-ch19
发表时间:
2013
期刊:
Pediatric ICD-10-CM 2018: A Manual for Provider-Based Coding
影响因子:
--
作者:
E. Aitini
通讯作者:
E. Aitini
影响因子:
7.2
作者:
E. Steyerberg
通讯作者:
E. Steyerberg
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
Morimune-Moriya Seira;Obara Keiya;Fuseya Marika;Katanosaka Masashi;岩田健史,関岳人,石川明日香,田村隆治,幾原雄一,柴田直哉
通讯作者:
岩田健史,関岳人,石川明日香,田村隆治,幾原雄一,柴田直哉
影响因子:
120.7
作者:
Kansagara, Devan;Englander, Honora;Salanitro, Amanda;Kagen, David;Theobald, Cecelia;Freeman, Michele;Kripalani, Sunil
通讯作者:
Kripalani, Sunil
DOI:
10.17632/bd63njzyjf.1
发表时间:
2021
期刊:
--
影响因子:
--
作者:
J. Adalsteinsson
通讯作者:
J. Adalsteinsson