Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks

Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks
复制标题

DOI:
--
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Zhishuai Guo;Mingrui Liu;Zhuoning Yuan;Li Shen;Wei Liu;Tianbao Yang
Zhishuai Guo;Mingrui Liu;Zhuoning Yuan;Li Shen;Wei Liu;Tianbao Yang
中科院分区:
其他
文献类型:
--
作者:
Zhishuai Guo;Mingrui Liu;Zhuoning Yuan;Li Shen;Wei Liu;Tianbao Yang

文献摘要

相似文献

在本文中,我们研究了使用深度神经网络作为预测模型的大规模AUC最大化的分布式算法。尽管分布式学习技术在深度学习中得到了广泛的研究,但由于其与标准损失最小化问题(例如,交叉熵)。为了解决这一挑战,我们提出并分析了一个通信高效的分布式优化算法的基础上的AUC最大化,其中每个工人和参数服务器之间的原始变量和对偶变量的通信只发生在每个工人的多个步骤的基于梯度的更新。与现有算法的朴素并行版本相比,该算法在单个机器上计算随机梯度并将其平均以更新模型参数,我们的算法需要的通信轮数少得多,理论上仍然实现了线性加速。据我们所知,这是\textbf{first}的工作,它以通信高效的分布式方式解决了深度神经网络的AUC最大化问题,同时仍然保持了理论上的线性加速特性。我们在几个基准数据集上的实验表明了我们的算法的有效性,也证实了我们的理论。
In this paper, we study distributed algorithms for large-scale AUC maximization with a deep neural network as a predictive model. Although distributed learning techniques have been investigated extensively in deep learning, they are not directly applicable to stochastic AUC maximization with deep neural networks due to its striking differences from standard loss minimization problems (e.g., cross-entropy). Towards addressing this challenge, we propose and analyze a communication-efficient distributed optimization algorithm based on a {\it non-convex concave} reformulation of the AUC maximization, in which the communication of both the primal variable and the dual variable between each worker and the parameter server only occurs after multiple steps of gradient-based updates in each worker. Compared with the naive parallel version of an existing algorithm that computes stochastic gradients at individual machines and averages them for updating the model parameter, our algorithm requires a much less number of communication rounds and still achieves a linear speedup in theory. To the best of our knowledge, this is the \textbf{first} work that solves the {\it non-convex concave min-max} problem for AUC maximization with deep neural networks in a communication-efficient distributed manner while still maintaining the linear speedup property in theory. Our experiments on several benchmark datasets show the effectiveness of our algorithm and also confirm our theory.