Implications of non-stationarity on predictive modeling using EHRs.

Implications of non-stationarity on predictive modeling using EHRs.
复制标题

DOI:
10.1016/j.jbi.2015.10.006
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Shah NH
Shah NH
中科院分区:
医学3区
文献类型:
--
作者:
Jung K;Shah NH

文献摘要

被引文献

相似文献

在电子健康记录(EHR)中捕获的临床信息量的快速增加已经导致越来越复杂的模型用于诸如疾病亚型发现和预测建模的目的的应用。然而,越来越多地采用EHR意味着在不久的将来,可用于此类目的的大部分数据将来自一个时期,在此期间,由于技术和激励措施的历史变化,医学实践和EHR的临床使用都在不断变化。在这项工作中,我们探讨了这种现象的影响,称为非平稳性,预测建模。我们专注于使用门诊伤口护理中心护理第一周期间EHR中可用的数据预测伤口延迟愈合的问题,使用了一个涵盖超过150,000个伤口和59,958名患者的大型数据集,为期四年。我们通过改变数据划分为训练集和测试集的方式来操纵模型开发过程中看到的非平稳性程度。我们证明,非平稳性可以导致相当不同的结论,不同的模型的预测能力和校准后验概率的相对优点。在该数据集中表现出的非平稳性下,复杂方法(如堆叠)相对于最佳简单分类器的性能优势消失了。因此,忽略非平稳性可能导致在此任务中的次优模型选择。
The rapidly increasing volume of clinical information captured in Electronic Health Records (EHRs) has led to the application of increasingly sophisticated models for purposes such as disease subtype discovery and predictive modeling. However, increasing adoption of EHRs implies that in the near future, much of the data available for such purposes will be from a time period during which both the practice of medicine and the clinical use of EHRs are in flux due to historic changes in both technology and incentives. In this work, we explore the implications of this phenomenon, called non-stationarity, on predictive modeling. We focus on the problem of predicting delayed wound healing using data available in the EHR during the first week of care in outpatient wound care centers, using a large dataset covering over 150,000 individual wounds and 59,958 patients seen over a period of four years. We manipulate the degree of non-stationarity seen by the model development process by changing the way data is split into training and test sets. We demonstrate that non-stationarity can lead to quite different conclusions regarding the relative merits of different models with respect to predictive power and calibration of their posterior probabilities. Under the non-stationarity exhibited in this dataset, the performance advantage of complex methods such as stacking relative to the best simple classifier disappears. Ignoring non-stationarity can thus lead to sub-optimal model selection in this task.