Identifiability of deep generative models without auxiliary information

Identifiability of deep generative models without auxiliary information
复制标题

DOI:
--
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Bohdan Kivva;Goutham Rajendran;Pradeep Ravikumar;Bryon Aragam
Bohdan Kivva;Goutham Rajendran;Pradeep Ravikumar;Bryon Aragam
中科院分区:
其他
文献类型:
--
作者:
Bohdan Kivva;Goutham Rajendran;Pradeep Ravikumar;Bryon Aragam

文献摘要

被引文献

相似文献

我们证明了一类广泛的深潜变量模型的可辨识性,这些模型(A)具有通用逼近能力,(B)是实际中常用的变分自动编码器的译码。与现有工作不同,我们的分析不需要弱监督、辅助信息或潜在空间的条件化。具体地说,我们证明了对于一类具有通用逼近能力的生成(即无监督)模型,边信息$u$是不必要的:我们证明了不观察$u$而只观察数据$x$的整个生成模型的可辨识性。我们考虑的模型与实践中使用的自动编码器架构相匹配,这些架构利用潜在空间中的混合先验和编码器中的RELU/泄漏RELU激活,例如VADE和MFC-VAE。我们的主要结果是一个可辨识性层次,它显著地概括了以前的工作,并揭示了不同的假设如何导致不同的可辨识性“强度”,并将具有各向同性高斯先验的某些“普通”VAE作为特例包括在内。例如,我们的最弱结果建立了直到仿射变换的(无监督)可辨识性,从而部分地解决了先前工作中提出的关于模型可辨识性的公开问题。通过对模拟数据和真实数据的实验验证了这些理论结果。
We prove identifiability of a broad class of deep latent variable models that (a) have universal approximation capabilities and (b) are the decoders of variational autoencoders that are commonly used in practice. Unlike existing work, our analysis does not require weak supervision, auxiliary information, or conditioning in the latent space. Specifically, we show that for a broad class of generative (i.e. unsupervised) models with universal approximation capabilities, the side information $u$ is not necessary: We prove identifiability of the entire generative model where we do not observe $u$ and only observe the data $x$. The models we consider match autoencoder architectures used in practice that leverage mixture priors in the latent space and ReLU/leaky-ReLU activations in the encoder, such as VaDE and MFC-VAE. Our main result is an identifiability hierarchy that significantly generalizes previous work and exposes how different assumptions lead to different"strengths"of identifiability, and includes certain"vanilla"VAEs with isotropic Gaussian priors as a special case. For example, our weakest result establishes (unsupervised) identifiability up to an affine transformation, and thus partially resolves an open problem regarding model identifiability raised in prior work. These theoretical results are augmented with experiments on both simulated and real data.