Understanding GANs in the LQG Setting: Formulation, Generalization and Stability

Understanding GANs in the LQG Setting: Formulation, Generalization and Stability
复制标题

DOI:
10.1109/jsait.2020.2991375
复制
发表时间:
2020-05
期刊:
IEEE Journal on Selected Areas in Information Theory
影响因子:
--
通讯作者:
S. Feizi;Farzan Farnia;Tony Ginart;David Tse
S. Feizi;Farzan Farnia;Tony Ginart;David Tse
中科院分区:
其他
文献类型:
--
作者:
S. Feizi;Farzan Farnia;Tony Ginart;David Tse

文献摘要

被引文献

相似文献

生成对抗网络(GANs)已经成为一种从数据中学习概率模型的流行方法。在本文中,我们在一个简单的LQG基准上提供了关于gan的基本问题的理解,包括它们的公式,泛化和稳定性,其中生成器是线性的,鉴别器是二次的,数据具有高维高斯分布。即使在这个简单的基准测试中,GAN问题也没有得到很好的理解,因为我们观察到现有的最先进的GAN架构可能由于(1)稳定性问题(即收敛到坏的局部解或根本不收敛)而无法学习适当的生成分布,(2)近似问题(即由不适当的GAN损失函数引起的不适当的全局GAN优化器),以及(3)泛化问题(即需要大量样本进行训练)。在这种设置中,我们提出了一种GAN架构,该架构可以恢复最大似然解并展示快速泛化。此外,我们分析了所提出GAN的不同计算方法的全局稳定性,并强调了它们的优缺点。最后,通过在MNIST和CIFAR-10数据集上的实验,我们概述了基于模型的方法的扩展,以在比考虑的高斯基准更复杂的设置中设计GAN。
Generative Adversarial Networks (GANs) have become a popular method to learn a probability model from data. In this paper, we provide an understanding of basic issues surrounding GANs including their formulation, generalization and stability on a simple LQG benchmark where the generator is Linear, the discriminator is Quadratic and the data has a high-dimensional Gaussian distribution. Even in this simple benchmark, the GAN problem has not been well-understood as we observe that existing state-of-the-art GAN architectures may fail to learn a proper generative distribution owing to (1) stability issues (i.e., convergence to bad local solutions or not converging at all), (2) approximation issues (i.e., having improper global GAN optimizers caused by inappropriate GAN’s loss functions), and (3) generalizability issues (i.e., requiring large number of samples for training). In this setup, we propose a GAN architecture which recovers the maximum-likelihood solution and demonstrates fast generalization. Moreover, we analyze global stability of different computational approaches for the proposed GAN and highlight their pros and cons. Finally, through experiments on MNIST and CIFAR-10 datasets, we outline extensions of our model-based approach to design GANs in more complex setups than the considered Gaussian benchmark.