Predicting 30-day Hospital Readmission with Publicly Available Administrative Database A Conditional Logistic Regression Modeling Approach

Predicting 30-day Hospital Readmission with Publicly Available Administrative Database A Conditional Logistic Regression Modeling Approach
复制标题

DOI:
10.3414/me14-02-0017
复制
发表时间:
2015-01-01
影响因子:
1.7
通讯作者:
Parikh, P.
Parikh, P.
中科院分区:
医学4区
文献类型:
--
作者:
Zhu, K.;Lou, Z.;Parikh, P.

文献摘要

被引文献

相似文献

简介:这篇文章是医学信息方法的焦点主题“医疗保健中的大数据和分析”的一部分。背景:医院再入院增加了医疗保健成本,并给医疗服务提供者和患者带来了巨大的痛苦。因此,医疗机构对预测哪些患者有再次入院的风险非常感兴趣。然而,目前基于Logistic回归的风险预测模型在应用于医院管理数据时具有有限的预测能力。同时,虽然决策树和随机森林已被应用,他们往往是太复杂的理解之间的医院practitioners.Objectives:探讨使用条件logistic回归,以提高预测准确性。方法:我们分析了HCUP全州范围内的住院病人出院记录数据集,其中包括病人的人口统计学,临床和护理利用数据从加州。我们提取了在11个月期间有住院经历的心力衰竭医疗保险受益人的记录。我们通过欠采样纠正了数据不平衡问题。在本研究中,我们首先应用标准逻辑回归与决策树来获得有影响的变数,并衍生出实际意义的决策规则。然后,我们对原始数据集进行相应的分层,并对每个数据层应用逻辑回归。我们进一步探讨了逻辑回归模型中交互变量的影响。我们进行了交叉验证,以评估条件Logistic回归(CSTR)的整体预测性能,并将其与标准分类模型进行比较。结果:开发的CSTR模型优于几个标准分类模型(例如,直接逻辑回归、逐步逻辑回归、随机森林、支持向量机)。例如,最好的logistic回归模型将分类准确率提高了近20%。此外,所开发的RISK模型倾向于实现比标准分类模型更好的10%以上的灵敏度,这可以被转化为加州州一年内额外400 - 500例心力衰竭患者再入院的正确标记。最后,确定从HCUP数据的几个关键的预测因素包括从放电的处置位置,慢性疾病的数量,和急性proceeds.Conclusions的数量:这将是有益的应用简单的决策规则,从决策树中获得一个特设的方式来指导队列分层。在建立不同数据层的logistic回归模型时,探索有影响力的预测因子之间的成对相互作用可能是有益的。明智地使用开发的ad-hoc预测模型为医院再入院预测模型的未来发展提供了见解,这可以在识别高风险患者和制定有效的出院后护理策略方面带来更好的直觉。最后,本文有望提高对收集其他标志物数据的认识,并为更大规模的再入院风险预测探索性研究开发必要的数据库基础设施。
Introduction: This article is part of the Focus Theme of Methods of Information in Medicine on "Big Data and Analytics in Healthcare".Background: Hospital readmissions raise healthcare costs and cause significant distress to providers and patients. It is, therefore, of great interest to healthcare organizations to predict what patients are at risk to be readmitted to their hospitals. However, current logistic regression based risk prediction models have limited prediction power when applied to hospital administrative data. Meanwhile, although decision trees and random forests have been applied, they tend to be too complex to understand among the hospital practitioners.Objectives: Explore the use of conditional logistic regression to increase the prediction accuracy.Methods: We analyzed an HCUP statewide inpatient discharge record dataset, which includes patient demographics, clinical and care utilization data from California. We extracted records of heart failure Medicare beneficiaries who had inpatient experience during an 11-month period. We corrected the data imbalance issue with under-sampling. In our study, we first applied standard logistic regression and decision tree to obtain influential variables and derive practically meaning decision rules. We then stratified the original data set accordingly and applied logistic regression on each data stratum. We further explored the effect of interacting variables in the logistic regression modeling. We conducted cross validation to assess the overall prediction performance of conditional logistic regression (CLR) and compared it with standard classification models.Results: The developed CLR models outperformed several standard classification models (e.g., straightforward logistic regression, stepwise logistic regression, random forest, support vector machine). For example, the best CLR model improved the classification accuracy by nearly 20% over the straightforward logistic regression model. Furthermore, the developed CLR models tend to achieve better sensitivity of more than 10% over the standard classification models, which can be translated to correct labeling of additional 400 - 500 readmissions for heart failure patients in the state of California over a year. Lastly, several key predictor identified from the HCUP data include the disposition location from discharge, the number of chronic conditions, and the number of acute procedures.Conclusions: It would be beneficial to apply simple decision rules obtained from the decision tree in an ad-hoc manner to guide the cohort stratification. It could be potentially beneficial to explore the effect of pairwise interactions between influential predictors when building the logistic regression models for different data strata. Judicious use of the ad-hoc CLR models developed offers insights into future development of prediction models for hospital readmissions, which can lead to better intuition in identifying high-risk patients and developing effective post-discharge care strategies. Lastly, this paper is expected to raise the awareness of collecting data on additional markers and developing necessary database infrastructure for larger-scale exploratory studies on readmission risk prediction.