Cascade Variational Auto-Encoder for Hierarchical Disentanglement

Cascade Variational Auto-Encoder for Hierarchical Disentanglement
复制标题

DOI:
10.1145/3511808.3557254
复制
发表时间:
2022-10
期刊:
Proceedings of the 31st ACM International Conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Fudong Lin;Xu Yuan;Lu Peng;N. Tzeng
Fudong Lin;Xu Yuan;Lu Peng;N. Tzeng
中科院分区:
其他
文献类型:
--
作者:
Fudong Lin;Xu Yuan;Lu Peng;N. Tzeng

文献摘要

被引文献

相似文献

虽然深度生成模型为许多新兴应用铺平了道路,但较大模型尺寸和复杂性的可解释性降低阻碍了它们对经济,安全,医疗保健等广泛领域的推广。考虑到这一障碍,一种常见的做法是通过潜在特征解纠缠来学习可解释的表示,旨在暴露一组相互独立的数据变化因素。然而,现有的方法要么无法捕捉的综合数据质量和模型的可解释性之间的权衡,或考虑的一阶功能解开,忽略了一个事实,即一个子集的显着特征可以携带可分解的语义含义,因此是高阶的性质。因此,我们在本文中提出了一种新的生成式建模范式,通过引入贝叶斯网络的正则化级联变分自动编码器(VAE)。具体来说,这个正则化器引导学习者发现一个表示空间,该表示空间包括一阶解纠缠特征和高阶显著特征,其中特征相互作用由贝叶斯结构捕获。实验表明,这种正则化给我们自由控制的表示空间,可以引导学习者发现可分解的语义之间的相互作用的独立因素。同时,我们在六个广泛使用的视觉数据集上进行了大量的基准实验,结果表明,我们的方法在合成数据质量和模型可解释性之间的权衡方面优于最先进的VAE竞争对手。虽然我们的设计是在VAE机制中构建的,但它实际上是通用的,并且可以更好地适应GAN和VAE,让它们同时享有高模型可解释性和高合成质量。
While deep generative models pave the way for many emerging applications, decreased interpretability for larger model sizes and complexities hinders their generalizability to wide domains such as economy, security, healthcare, etc. Considering this obstacle, a common practice is to learn interpretable representations through latent feature disentanglement, aiming for exposing a set of mutually independent factors of data variations. However, existing methods either fail to catch the trade-off between the synthetic data quality and model interpretability, or consider the first-order feature disentangling only, overlooking the fact that a subset of salient features can carry decomposable semantic meanings and hence be of high-order in nature. Hence, we in this paper propose a novel generative modeling paradigm by introducing a Bayesian network-based regularize on a cascade Variational Auto-Encoder (VAE). Specifically, this regularizer guides the learner to discover a representation space that comprises both first-order disentangled features and high-order salient features, with the feature interplay captured by the Bayesian structure. Experiments demonstrate that this regularizer gives us free control over the representation space and can guide the learner to discover decomposable semantic meanings by capturing the interplay among independent factors. Meanwhile, we benchmark extensive experiments on six widely-used vision datasets, and the results exhibit that our approach outperforms the state-of-the-art VAE competitors in terms of the trade-off between the synthetic data quality and model interpretability. Although our design is framed in the VAE regime, it in effect is generic and can be better amenable to both GANs and VAEs in terms of letting them concurrently enjoy both high model interpretability and high synthesis quality.