Adapting machine learning techniques to censored time-to-event health record data: A general-purpose approach using inverse probability of censoring weighting.

Adapting machine learning techniques to censored time-to-event health record data: A general-purpose approach using inverse probability of censoring weighting.
复制标题

DOI:
10.1016/j.jbi.2016.03.009
复制
发表时间:
2016-06
影响因子:
4.5
通讯作者:
O'Connor PJ
O'Connor PJ
中科院分区:
医学3区
文献类型:
--
作者:
Vock DM;Wolfson J;Bandyopadhyay S;Adomavicius G;Johnson PE;Vazquez-Benitez G;O'Connor PJ

文献摘要

被引文献

相似文献

基于患者个体特征预测在特定时间范围内经历各种健康结果或不良事件(例如,在接下来的5年内发生心脏病)的概率的模型是管理患者护理的重要工具。电子健康数据(EHD)是极具吸引力的培训数据来源,因为它们提供了对来自当今患者群体的大量丰富的个人级别数据的访问。然而,由于EHD是通过从行政和临床数据库中提取信息而得出的,因此在一个人想要做出预测的整个时间范围内,一些受试者将不会受到观察;这种对后续行动的丧失通常是由于从卫生系统中注销。对于没有完全随访的受试者,他们是否经历了不良事件是未知的,在统计学上,事件时间被认为是正确审查的。解决该问题的大多数机器学习方法都是相对特别的;例如,处理事件状态未知的观测的常见方法包括1)丢弃这些观测,2)将它们视为非事件,3)将这些观测分为两个观测:一个是事件发生的地方,另一个不是事件发生的地方。在这篇文章中,我们提出了一种通用的方法来解释右审查结果使用审查权重的逆概率(IPCW)。我们说明了如何轻松地将IPCW合并到许多现有的机器学习算法中,这些算法用于挖掘大型医疗保健数据,包括贝叶斯网络、k近邻、决策树和广义加法模型。然后,我们表明,当使用来自美国中西部大型医疗系统的EHD来预测经历心血管不良事件的5年风险时,我们的方法比三种特别方法导致更好的校准预测。
Models for predicting the probability of experiencing various health outcomes or adverse events over a certain time frame (e.g., having a heart attack in the next 5 years) based on individual patient characteristics are important tools for managing patient care. Electronic health data (EHD) are appealing sources of training data because they provide access to large amounts of rich individual-level data from present-day patient populations. However, because EHD are derived by extracting information from administrative and clinical databases, some fraction of subjects will not be under observation for the entire time frame over which one wants to make predictions; this loss to follow-up is often due to disenrollment from the health system. For subjects without complete follow-up, whether or not they experienced the adverse event is unknown, and in statistical terms the event time is said to be right-censored. Most machine learning approaches to the problem have been relatively ad hoc; for example, common approaches for handling observations in which the event status is unknown include 1) discarding those observations, 2) treating them as non-events, 3) splitting those observations into two observations: one where the event occurs and one where the event does not. In this paper, we present a general-purpose approach to account for right-censored outcomes using inverse probability of censoring weighting (IPCW). We illustrate how IPCW can easily be incorporated into a number of existing machine learning algorithms used to mine big health care data including Bayesian networks, k-nearest neighbors, decision trees, and generalized additive models. We then show that our approach leads to better calibrated predictions than the three ad hoc approaches when applied to predicting the 5-year risk of experiencing a cardiovascular adverse event, using EHD from a large U.S. Midwestern healthcare system.