Learning to Generate with Memory

Learning to Generate with Memory
复制标题

DOI:
--
复制
发表时间:
2016-02
期刊:
--
影响因子:
--
通讯作者:
Chongxuan Li;Jun Zhu;Bo Zhang-
Chongxuan Li;Jun Zhu;Bo Zhang-
中科院分区:
其他
文献类型:
--
作者:
Chongxuan Li;Jun Zhu;Bo Zhang-

文献摘要

被引文献

相似文献

记忆单元已被广泛用于丰富深度网络捕获推理和预测任务中的长期依赖性的能力,但对擅长从未标记数据推断高级不变表示的深度生成模型(DGM)的研究很少。本文提出了一种深度生成模型,具有可能较大的外部存储器和注意机制,以捕获表示学习中自下而上的抽象过程中经常丢失的局部细节信息。通过采用平滑注意力模型,通过自动编码变分贝叶斯方法优化数据似然的变分界限,对整个网络进行端到端训练,其中联合学习不对称识别网络以推断高级不变表示。非对称架构可以减少自下而上的不变特征提取和自上而下的实例细节生成之间的竞争。我们在多个数据集上的实验表明,内存可以显着提高 DGM 在各种任务上的性能,包括密度估计、图像生成和缺失值插补,并且具有内存的 DGM 可以实现最先进的定量结果。
Memory units have been widely used to enrich the capabilities of deep networks on capturing long-term dependencies in reasoning and prediction tasks, but little investigation exists on deep generative models (DGMs) which are good at inferring high-level invariant representations from unlabeled data. This paper presents a deep generative model with a possibly large external memory and an attention mechanism to capture the local detail information that is often lost in the bottom-up abstraction process in representation learning. By adopting a smooth attention model, the whole network is trained end-to-end by optimizing a variational bound of data likelihood via auto-encoding variational Bayesian methods, where an asymmetric recognition network is learnt jointly to infer high-level invariant representations. The asymmetric architecture can reduce the competition between bottom-up invariant feature extraction and top-down generation of instance details. Our experiments on several datasets demonstrate that memory can significantly boost the performance of DGMs on various tasks, including density estimation, image generation, and missing value imputation, and DGMs with memory can achieve state-of-the-art quantitative results.