FL-NTK: A Neural Tangent Kernel-based Framework for Federated Learning Convergence Analysis

FL-NTK: A Neural Tangent Kernel-based Framework for Federated Learning Convergence Analysis
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Baihe Huang;Xiaoxiao Li;Zhao Song;Xin Yang
Baihe Huang;Xiaoxiao Li;Zhao Song;Xin Yang
中科院分区:
其他
文献类型:
--
作者:
Baihe Huang;Xiaoxiao Li;Zhao Song;Xin Yang

文献摘要

相似文献

联合学习(FL)是一种新兴的学习方案,它允许不同的分布式客户端在不共享数据的情况下一起训练深度神经网络。神经网络因其前所未有的成功而变得流行起来。就我们所知,关于具有显式形式和多步更新的神经网络的FL的理论保证还没有被探索。然而,神经网络在FL中的训练分析是不平凡的,原因有两个:第一,我们正在优化的目标损失函数是非光滑和非凸的,第二,我们甚至没有在梯度方向上进行更新。现有的基于梯度下降方法的收敛结果在很大程度上依赖于梯度方向用于更新这一事实。受神经切核(NTK)分析的启发,提出了一种新的FL收敛分析方法--联邦学习神经切核(FL-NTK),它对应于FL中通过梯度下降训练的过参数RELU神经网络。理论上,FL-NTK在适当调整学习参数的情况下,以线性速度收敛到全局最优解。此外,在适当的分布假设下,FL-NTK也能达到较好的泛化效果。
Federated Learning (FL) is an emerging learning scheme that allows different distributed clients to train deep neural networks together without data sharing. Neural networks have become popular due to their unprecedented success. To the best of our knowledge, the theoretical guarantees of FL concerning neural networks with explicit forms and multi-step updates are unexplored. Nevertheless, training analysis of neural networks in FL is non-trivial for two reasons: first, the objective loss function we are optimizing is non-smooth and non-convex, and second, we are even not updating in the gradient direction. Existing convergence results for gradient descent-based methods heavily rely on the fact that the gradient direction is used for updating. This paper presents a new class of convergence analysis for FL, Federated Learning Neural Tangent Kernel (FL-NTK), which corresponds to overparamterized ReLU neural networks trained by gradient descent in FL and is inspired by the analysis in Neural Tangent Kernel (NTK). Theoretically, FL-NTK converges to a global-optimal solution at a linear rate with properly tuned learning parameters. Furthermore, with proper distributional assumptions, FL-NTK can also achieve good generalization.