Handling incomplete heterogeneous data using VAEs

Handling incomplete heterogeneous data using VAEs
复制标题

DOI:
10.1016/j.patcog.2020.107501
复制
发表时间:
2020-11-01
影响因子:
8
通讯作者:
Valera, Isabel
Valera, Isabel
中科院分区:
计算机科学1区
文献类型:
--
作者:
Nazabal, Alfredo;Olmos, Pablo M.;Valera, Isabel

文献摘要

被引文献

相似文献

变分自动编码器(VAE)以及其他生成模型已被证明可以有效且准确地捕获大量复杂高维数据的潜在结构。然而,现有的 VAE 仍然无法直接处理异构(混合连续和离散)或不完整(随机丢失数据)的数据,这在实际应用中确实很常见。在本文中,我们提出了一个通用框架来设计适合拟合不完整异构数据的 VAE。所提出的 HI-VAE 包括实值、正实值、区间、分类、序数和计数数据的似然模型,并允许对缺失数据进行准确估计(以及可能的插补)。此外,HI-VAE 在监督任务中呈现出有竞争力的预测性能,在对不完整数据进行训练时优于监督模型。 (C) 2020 Elsevier Ltd. 保留所有权利。
Variational autoencoders (VAEs), as well as other generative models, have been shown to be efficient and accurate for capturing the latent structure of vast amounts of complex high-dimensional data. However, existing VAEs can still not directly handle data that are heterogenous (mixed continuous and discrete) or incomplete (with missing data at random), which is indeed common in real-world applications.In this paper, we propose a general framework to design VAEs suitable for fitting incomplete heterogenous data. The proposed HI-VAE includes likelihood models for real-valued, positive real valued, interval, categorical, ordinal and count data, and allows accurate estimation (and potentially imputation) of missing data. Furthermore, HI-VAE presents competitive predictive performance in supervised tasks, outperforming supervised models when trained on incomplete data. (C) 2020 Elsevier Ltd. All rights reserved.