Data Pre-Processing Using Neural Processes for Modeling Personalized Vital-Sign Time-Series Data

Data Pre-Processing Using Neural Processes for Modeling Personalized Vital-Sign Time-Series Data
复制标题

使用神经过程进行数据预处理,对个性化生命体征时间序列数据进行建模

DOI:
10.1109/jbhi.2021.3107518
复制
发表时间:
2021
影响因子:
7.7
通讯作者:
D. Clifton
D. Clifton
中科院分区:
工程技术1区
文献类型:
--
作者:
Pulkit Sharma;Farah E. Shamout;V. Abrol;D. Clifton

文献摘要

被引文献

相似文献

从电子病历中检索的临床时间序列数据被广泛用于构建不良事件的预测模型,以支持资源管理。这些数据通常是稀疏和不规则采样的,这使得使用许多常见的机器学习方法具有挑战性。缺失值可以通过向前推进最后一个值或通过线性回归进行插值。高斯过程(GP)回归也用于执行插补,并且通常以规则的间隔对时间序列进行重新采样。GP的使用可能需要广泛的,可能是临时的调查,以确定模型结构,如适当的协方差函数。这对于多变量真实世界的临床数据可能具有挑战性,其中时间序列变量彼此表现出不同的动态。在这项工作中,我们构建生成模型,使用神经隐变量模型(称为神经过程(NP))来估计临床时间序列数据中的缺失值。NP模型在潜在空间中采用条件先验分布,通过对局部水平的变化进行建模来学习数据中的全局不确定性。与传统的生成式建模相比,这种先验并不是固定的,而是在训练过程中学习的。因此,NP模型提供了适应可用临床数据动态的灵活性。我们提出了一个NP框架的变体,用于对潜在空间和输入空间之间的互信息进行有效建模,确保有意义的先验知识。使用MIMIC III数据集的实验表明,所提出的方法相比,传统的方法的有效性。
Clinical time-series data retrieved from electronic medical records are widely used to build predictive models of adverse events to support resource management. Such data is often sparse and irregularly-sampled, which makes it challenging to use many common machine learning methods. Missing values may be interpolated by carrying the last value forward, or through linear regression. Gaussian process (GP) regression is also used for performing imputation, and often re-sampling of time-series at regular intervals. The use of GPs can require extensive, and likely adhoc, investigation to determine model structure, such as an appropriate covariance function. This can be challenging for multivariate real-world clinical data, in which time-series variables exhibit different dynamics to one another. In this work, we construct generative models to estimate missing values in clinical time-series data using a neural latent variable model, known as a Neural Process (NP). The NP model employs a conditional prior distribution in the latent space to learn global uncertainty in the data by modelling variations at a local level. In contrast to conventional generative modelling, this prior is not fixed and is itself learned during the training process. Thus, NP model provides the flexibility to adapt to the dynamics of the available clinical data. We propose a variant of the NP framework for efficient modelling of the mutual information between the latent and input spaces, ensuring meaningful learned priors. Experiments using the MIMIC III dataset demonstrate the effectiveness of the proposed approach as compared to conventional methods.