Deconstructing Generative Adversarial Networks

Deconstructing Generative Adversarial Networks
复制标题

DOI:
10.1109/tit.2020.2983698
复制
发表时间:
2019-01
影响因子:
2.5
通讯作者:
Banghua Zhu;Jiantao Jiao;David Tse
Banghua Zhu;Jiantao Jiao;David Tse
中科院分区:
计算机科学2区
文献类型:
--
作者:
Banghua Zhu;Jiantao Jiao;David Tse

文献摘要

被引文献

相似文献

生成性对抗性网络(GANS)是一种蓬勃发展的无监督机器学习技术,它已经在计算机视觉、自然语言处理等各个领域取得了重大进展。然而,众所周知,GAN很难训练,而且通常会遭受模式崩溃和鉴别器获胜的问题。为了解释甘斯的经验观察,并设计出更好的经验,我们将甘斯的研究解构为三个组成部分,并做出了以下贡献。·公式:我们提出了Gans人口目标的微扰观。基于这种解释,我们证明了GANS可以连接到稳健统计框架,并提出了一种新的GAN结构,称为级联GANS,当真实分布是高维的且被离群值破坏时,可以证明地恢复有意义的低维生成器近似。·推广:给定GANS的总体目标,我们设计了一个系统的原则,在允许距离下,设计GANS来设计仅使用有限样本来满足总体要求的GAN。我们在三种情况下实现了我们的原理,以实现多项式甚至有时是接近最优的样本复杂性:(1)在任意伪范数下学习任意生成元;(2)在全变差距离下学习高斯位置族,其中我们利用我们的原理为被视为Gans的Tukey中值的近最优性提供一个新的证明;(3)在Wasserstein距离下学习高维任意分布的低维高斯近似。我们演示了GANS中近似误差和统计误差之间的基本权衡,并演示了如何将我们的原理应用于实践,仅使用经验样本来预测GAN的样本数量,以便不遭受鉴别器获胜问题。·优化:我们证明了交替梯度下降在优化GAN主元分析公式时不是局部渐近稳定的。我们发现最小极大对偶间隙非零可能是原因之一,并提出了一种新的GAN结构,其对偶间隙为零,其中博弈的值等于先前的极小极大值(而不是最大值)。我们证明了新的GAN结构在交替梯度下降下解PCA时是全局渐近稳定的。
Generative Adversarial Networks (GANs) are a thriving unsupervised machine learning technique that has led to significant advances in various fields such as computer vision, natural language processing, among others. However, GANs are known to be difficult to train and usually suffer from mode collapse and the discriminator winning problem. To interpret the empirical observations of GANs and design better ones, we deconstruct the study of GANs into three components and make the following contributions. •Formulation: we propose a perturbation view of the population target of GANs. Building on this interpretation, we show that GANs can be connected to the robust statistics framework, and propose a novel GAN architecture, termed as Cascade GANs, to provably recover meaningful low-dimensional generator approximations when the real distribution is high-dimensional and corrupted by outliers.•Generalization: given a population target of GANs, we design a systematic principle, projection under admissible distance, to design GANs to meet the population requirement using only finite samples. We implement our principle in three cases to achieve polynomial and sometimes near-optimal sample complexities: (1) learning an arbitrary generator under an arbitrary pseudonorm; (2) learning a Gaussian location family under total variation distance, where we utilize our principle to provide a new proof for the near-optimality of the Tukey median viewed as GANs; (3) learning a low-dimensional Gaussian approximation of a high-dimensional arbitrary distribution under Wasserstein distance. We demonstrate a fundamental trade-off in the approximation error and statistical error in GANs, and demonstrate how to apply our principle in practice with only empirical samples to predict how many samples would be sufficient for GANs in order not to suffer from the discriminator winning problem.•Optimization: we demonstrate alternating gradient descent is provably not locally asymptotically stable in optimizing the GAN formulation of PCA. We found that the minimax duality gap being non-zero might be one of the causes, and propose a new GAN architecture whose duality gap is zero, where the value of the game is equal to the previous minimax value (not the maximin value). We prove the new GAN architecture is globally asymptotically stable in solving PCA under alternating gradient descent.