A Decentralized Parallel Algorithm for Training Generative Adversarial Nets

A Decentralized Parallel Algorithm for Training Generative Adversarial Nets
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
arXiv: Optimization and Control
影响因子:
--
通讯作者:
Mingrui Liu;Wei Zhang;Youssef Mroueh;Xiaodong Cui;Jarret Ross;Tianbao Yang;Payel Das
Mingrui Liu;Wei Zhang;Youssef Mroueh;Xiaodong Cui;Jarret Ross;Tianbao Yang;Payel Das
中科院分区:
其他
文献类型:
--
作者:
Mingrui Liu;Wei Zhang;Youssef Mroueh;Xiaodong Cui;Jarret Ross;Tianbao Yang;Payel Das

文献摘要

相似文献

生成对抗网络(GAN)是深度学习社区中强大的一类生成模型。目前大规模GAN训练的实践~\citep{brock2018large}利用大型模型和分布式大批量训练策略,并在集中式设计的深度学习框架(例如TensorFlow、PyTorch等)上实现。在中心化的网络拓扑中,每个worker都需要与中心节点进行通信。然而,当网络带宽较低或网络延迟较高时,性能会显着下降。尽管最近在训练深度神经网络的去中心化算法方面取得了进展,但仍不清楚是否可以以去中心化的方式训练 GAN。主要困难在于同时处理非凸非凹最小最大优化和去中心化通信。在本文中,我们通过设计\textbf{第一个基于梯度的去中心化并行算法}来解决这个难题,该算法允许工作人员在一次迭代中进行多轮通信,并同时更新鉴别器和生成器,并且这种设计使其适合所提出的去中心化算法的收敛分析。理论上,我们提出的分散算法能够解决一类非凸非凹最小-最大问题,并可证明非渐近收敛到一阶驻点。 GAN 上的实验结果证明了所提出算法的有效性。
Generative Adversarial Networks (GANs) are powerful class of generative models in the deep learning community. Current practice on large-scale GAN training~\citep{brock2018large} utilizes large models and distributed large-batch training strategies, and is implemented on deep learning frameworks (e.g., TensorFlow, PyTorch, etc.) designed in a centralized manner. In the centralized network topology, every worker needs to communicate with the central node. However, when the network bandwidth is low or network latency is high, the performance would be significantly degraded. Despite recent progress on decentralized algorithms for training deep neural networks, it remains unclear whether it is possible to train GANs in a decentralized manner. The main difficulty lies at handling the nonconvex-nonconcave min-max optimization and the decentralized communication simultaneously. In this paper, we address this difficulty by designing the \textbf{first gradient-based decentralized parallel algorithm} which allows workers to have multiple rounds of communications in one iteration and to update the discriminator and generator simultaneously, and this design makes it amenable for the convergence analysis of the proposed decentralized algorithm. Theoretically, our proposed decentralized algorithm is able to solve a class of non-convex non-concave min-max problems with provable non-asymptotic convergence to first-order stationary point. Experimental results on GANs demonstrate the effectiveness of the proposed algorithm.