Variational Autoencoders for Sparse and Overdispersed Discrete Data

Variational Autoencoders for Sparse and Overdispersed Discrete Data
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
He Zhao;Piyush Rai;Lan Du;Wray L. Buntine;Mingyuan Zhou
He Zhao;Piyush Rai;Lan Du;Wray L. Buntine;Mingyuan Zhou
中科院分区:
其他
文献类型:
--
作者:
He Zhao;Piyush Rai;Lan Du;Wray L. Buntine;Mingyuan Zhou

文献摘要

相似文献

许多应用程序,如文本建模,高通量测序和推荐系统,需要分析稀疏,高维和过度分散的离散(计数值或二进制)数据。虽然概率矩阵分解和线性/非线性潜在因子模型在模拟这些数据方面取得了巨大的成功,但由于在计数值数据和模型错误指定中建模过度分散的能力不足,许多现有模型可能具有较差的建模性能。在本文中,我们全面研究了这些问题,并提出了一个变分自动编码器的框架,通过负二项分布生成离散数据。我们还检查了模型捕获特性的能力,例如离散数据中的自激和交叉激励,这对于过度分散建模至关重要。我们对离散数据分析中的三个重要问题进行了广泛的实验:文本分析,协同过滤和多标签学习。与几个国家的最先进的基线相比,所提出的模型在上述问题上取得了显着更好的性能。
Many applications, such as text modelling, high-throughput sequencing, and recommender systems, require analysing sparse, high-dimensional, and overdispersed discrete (count-valued or binary) data. Although probabilistic matrix factorisation and linear/nonlinear latent factor models have enjoyed great success in modelling such data, many existing models may have inferior modelling performance due to the insufficient capability of modelling overdispersion in count-valued data and model misspecification in general. In this paper, we comprehensively study these issues and propose a variational autoencoder based framework that generates discrete data via negative-binomial distribution. We also examine the model's ability to capture properties, such as self- and cross-excitations in discrete data, which is critical for modelling overdispersion. We conduct extensive experiments on three important problems from discrete data analysis: text analysis, collaborative filtering, and multi-label learning. Compared with several state-of-the-art baselines, the proposed models achieve significantly better performance on the above problems.