Identifiable Deep Generative Models via Sparse Decoding

Identifiable Deep Generative Models via Sparse Decoding
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
Trans. Mach. Learn. Res.
影响因子:
--
通讯作者:
Gemma E. Moran;Dhanya Sridhar;Yixin Wang;D. Blei
Gemma E. Moran;Dhanya Sridhar;Yixin Wang;D. Blei
中科院分区:
其他
文献类型:
--
作者:
Gemma E. Moran;Dhanya Sridhar;Yixin Wang;D. Blei

文献摘要

被引文献

相似文献

我们开发了用于高维数据无监督表示学习的稀疏VAE。稀疏VAE学习一组潜在因素(表示),这些潜在因素(表示)总结了观察到的数据特征中的关联。底层模型是稀疏的,因为每个观察到的特征(即数据的每个维度)依赖于潜在因素的一个小子集。举个例子,在评分数据中,每部电影只被几个类型所描述;在文本数据中,每个词只适用于少数几个主题;在基因组学中,每个基因只在少数生物过程中活跃。我们证明了这种稀疏深度生成模型是可识别的:在无限数据下,可以学习到真实的模型参数。(相比之下,大多数深度生成模型都是不可识别的。)我们用模拟数据和真实数据对稀疏VAE进行了实证研究。我们发现,该方法恢复了有意义的潜在因素,并且比相关方法具有更小的遗漏重建误差。
We develop the sparse VAE for unsupervised representation learning on high-dimensional data. The sparse VAE learns a set of latent factors (representations) which summarize the associations in the observed data features. The underlying model is sparse in that each observed feature (i.e. each dimension of the data) depends on a small subset of the latent factors. As examples, in ratings data each movie is only described by a few genres; in text data each word is only applicable to a few topics; in genomics, each gene is active in only a few biological processes. We prove such sparse deep generative models are identifiable: with infinite data, the true model parameters can be learned. (In contrast, most deep generative models are not identifiable.) We empirically study the sparse VAE with both simulated and real data. We find that it recovers meaningful latent factors and has smaller heldout reconstruction error than related methods.