AE-OT-GAN: Training GANs from data specific latent distribution

AE-OT-GAN: Training GANs from data specific latent distribution
复制标题

DOI:
10.1007/978-3-030-58574-7_33
复制
发表时间:
2020-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Dongsheng An;Yang Guo;Min Zhang;Xin Qi;Na Lei;S. Yau;X. Gu
Dongsheng An;Yang Guo;Min Zhang;Xin Qi;Na Lei;S. Yau;X. Gu
中科院分区:
其他
文献类型:
--
作者:
Dongsheng An;Yang Guo;Min Zhang;Xin Qi;Na Lei;S. Yau;X. Gu

文献摘要

相似文献

虽然生成式对抗网络(GANs)是生成逼真、清晰图像的重要模型,但其训练不稳定且存在模态崩溃问题。gan的问题来自于用连续dnn逼近内在的不连续分布变换映射。最近提出的AE-OT模型通过显式计算自编码器潜在空间中的不连续最优变换映射来解决不连续问题。AE-OT生成的图像虽然没有模式坍缩,但是图像模糊。在本文中,我们提出了AE-OT-GAN模型,利用这两种模型的优点:生成高质量的图像,同时克服了模态崩溃问题。具体来说,我们首先通过自编码器(AE)将低维图像流形嵌入到潜在空间中。然后利用扩展的半离散最优传输(SDOT)映射生成新的潜码。最后,我们的GAN模型经过训练,可以从扩展的SDOT地图引起的潜在分布中生成高质量的图像。从该数据集相关的潜在分布到数据分布的分布变换映射将是连续的,因此可以很好地近似连续dnn。此外,潜在代码和真实图像之间的配对数据给我们提供了对生成器的进一步限制,并稳定了训练过程。在简单MNIST数据集和CIFAR10、CelebA等复杂数据集上的实验表明了该方法的优越性。
Though generative adversarial networks (GANs) are prominent models to generate realistic and crisp images, they are unstable to train and suffer from the mode collapse problem. The problems of GANs come from approximating the intrinsic discontinuous distribution transform map with continuous DNNs. The recently proposed AE-OT model addresses the discontinuity problem by explicitly computing the discontinuous optimal transform map in the latent space of the autoencoder. Though have no mode collapse, the generated images by AE-OT are blurry. In this paper, we propose the AE-OT-GAN model to utilize the advantages of the both models: generate high quality images and at the same time overcome the mode collapse problems. Specifically, we firstly embed the low dimensional image manifold into the latent space by autoencoder (AE). Then the extended semi-discrete optimal transport (SDOT) map is used to generate new latent codes. Finally, our GAN model is trained to generate high quality images from the latent distribution induced by the extended SDOT map. The distribution transform map from this dataset related latent distribution to the data distribution will be continuous, and thus can be well approximated by the continuous DNNs. Additionally, the paired data between the latent codes and the real images gives us further restriction about the generator and stabilizes the training process. Experiments on simple MNIST dataset and complex datasets like CIFAR10 and CelebA show the advantages of the proposed method.