GANs as Gradient Flows that Converge

GANs as Gradient Flows that Converge
复制标题

DOI:
10.48550/arxiv.2205.02910
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yu‐Jui Huang;Yuchong Zhang
Yu‐Jui Huang;Yuchong Zhang
中科院分区:
其他
文献类型:
--
作者:
Yu‐Jui Huang;Yuchong Zhang

文献摘要

相似文献

本文在概率密度函数空间中用梯度下降法研究无监督学习问题。主要结果表明,沿着由分布相关的常微分方程(ODE)诱导的梯度流,未知数据分布以长时间极限的形式出现。也就是说,我们可以通过模拟依赖于分布的常微分方程来揭示数据的分布。有趣的是,ODE的模拟相当于生成对抗网络(GAN)的训练。这种等价性为GAN提供了一种新的“合作”观点,更重要的是,它为GAN的分歧提供了新的视角。特别是,它揭示了GAN算法隐式地最小化了两组样本之间的均方误差(MSE),而这种MSE拟合本身就可能导致GAN发散。为了构造分布依赖常微分方程的解,我们首先利用Banach空间中微分方程的Crandall-Liggett定理证明了相应的非线性Fokker-Planck方程存在唯一的弱解.在此基础上,利用Trevisan叠加原理,构造了一个唯一的解。通过对Fokker-Planck方程的分析,得到了诱导梯度流对数据分布的收敛性。
This paper approaches the unsupervised learning problem by gradient descent in the space of probability density functions. A main result shows that along the gradient flow induced by a distribution-dependent ordinary differential equation (ODE), the unknown data distribution emerges as the long-time limit. That is, one can uncover the data distribution by simulating the distribution-dependent ODE. Intriguingly, the simulation of the ODE is shown equivalent to the training of generative adversarial networks (GANs). This equivalence provides a new"cooperative"view of GANs and, more importantly, sheds new light on the divergence of GANs. In particular, it reveals that the GAN algorithm implicitly minimizes the mean squared error (MSE) between two sets of samples, and this MSE fitting alone can cause GANs to diverge. To construct a solution to the distribution-dependent ODE, we first show that the associated nonlinear Fokker-Planck equation has a unique weak solution, by the Crandall-Liggett theorem for differential equations in Banach spaces. Based on this solution to the Fokker-Planck equation, we construct a unique solution to the ODE, using Trevisan's superposition principle. The convergence of the induced gradient flow to the data distribution is obtained by analyzing the Fokker-Planck equation.