A Gaussian Latent Variable Model for Incomplete Mixed Type Data

A Gaussian Latent Variable Model for Incomplete Mixed Type Data
复制标题

DOI:
10.1109/icassp49357.2023.10095772
复制
发表时间:
2023-06
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Marzieh Ajirak;P. Djurić
Marzieh Ajirak;P. Djurić
中科院分区:
其他
文献类型:
--
作者:
Marzieh Ajirak;P. Djurić

文献摘要

相似文献

在许多机器学习问题中,必须处理不同类型的数据,包括连续、离散和分类数据。此外,通常情况下,这些数据中的许多数据从数据库中丢失。本文提出了一个高斯过程框架,有效地捕捉信息的混合数值和分类数据,有效地纳入缺失的变量。首先,我们提出了一个混合类型数据的生成模型。生成模型利用高斯过程,其内核从潜在向量构建。我们还提出了一种方法的推理的未知数,并在其实施中,我们依赖于稀疏谱近似的高斯过程和变分推理。我们证明了该方法的监督和无监督任务的性能。首先,我们研究了在无监督环境中缺失变量的插补,然后我们展示了IBM员工数据的联合插补和分类结果。
In many machine learning problems, one has to work with data of different types, including continuous, discrete, and categorical data. Further, it is often the case that many of these data are missing from the database. This paper proposes a Gaussian process framework that efficiently captures the information from mixed numerical and categorical data that effectively incorporates missing variables. First, we propose a generative model for the mixed-type data. The generative model exploits Gaussian processes with kernels constructed from the latent vectors. We also propose a method for inference of the unknowns, and in its implementation, we rely on a sparse spectrum approximation of the Gaussian processes and variational inference. We demonstrate the performance of the method for both supervised and unsupervised tasks. First, we investigate the imputation of missing variables in an unsupervised setting, and then we show the results of joint imputation and classification on IBM employee data.