VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data

VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Chao Ma-;Sebastian Tschiatschek;José Miguel Hernández-Lobato;Richard E. Turner;Cheng Zhang
Chao Ma-;Sebastian Tschiatschek;José Miguel Hernández-Lobato;Richard E. Turner;Cheng Zhang
中科院分区:
其他
文献类型:
--
作者:
Chao Ma-;Sebastian Tschiatschek;José Miguel Hernández-Lobato;Richard E. Turner;Cheng Zhang

文献摘要

被引文献

相似文献

由于自然数据集的异质性,深度生成模型在现实世界的应用中通常表现不佳。异质性来自于包含不同类型特征的数据(分类、有序、连续等)。并且相同类型的特征具有不同的边缘分布。我们提出了一个扩展的变分自动编码器(VAE)称为VAEM处理这样的异构数据。VAEM是一种深度生成模型,它以两阶段的方式进行训练,使得第一阶段为第二阶段提供更统一的数据表示,从而避免了异构数据引起的问题。我们提供了VAEM的扩展来处理部分观测数据,并展示了其在数据生成,丢失数据预测和顺序特征选择任务中的性能。我们的研究结果表明,VAEM拓宽了可以成功部署深度生成模型的现实世界应用范围。
Deep generative models often perform poorly in real-world applications due to the heterogeneity of natural data sets. Heterogeneity arises from data containing different types of features (categorical, ordinal, continuous, etc.) and features of the same type having different marginal distributions. We propose an extension of variational autoencoders (VAEs) called VAEM to handle such heterogeneous data. VAEM is a deep generative model that is trained in a two stage manner such that the first stage provides a more uniform representation of the data to the second stage, thereby sidestepping the problems caused by heterogeneous data. We provide extensions of VAEM to handle partially observed data, and demonstrate its performance in data generation, missing data prediction and sequential feature selection tasks. Our results show that VAEM broadens the range of real-world applications where deep generative models can be successfully deployed.