Decentralized and Model-Free Federated Learning: Consensus-Based Distillation in Function Space

Decentralized and Model-Free Federated Learning: Consensus-Based Distillation in Function Space
复制标题

DOI:
10.1109/tsipn.2022.3205549
复制
发表时间:
2021-04
影响因子:
3.2
通讯作者:
Akihito Taya;T. Nishio;M. Morikura;Koji Yamamoto
Akihito Taya;T. Nishio;M. Morikura;Koji Yamamoto
中科院分区:
计算机科学2区
文献类型:
--
作者:
Akihito Taya;T. Nishio;M. Morikura;Koji Yamamoto

文献摘要

相似文献

针对通过多跳网络连接的万物互联设备,提出了一种完全去中心化的联合学习(FL)方案。由于FL算法很难收敛于机器学习(ML)模型的参数,本文主要研究ML模型在函数空间中的收敛问题。考虑到最大似然任务的代表性损失函数,如均方误差(MSE)和Kullback-Leibler(KL)发散度,都是凸泛函,直接更新函数空间中的函数的算法可以收敛到最优解。本文的核心概念是定制一种基于共识的优化算法,使其工作在函数空间,并以分布式的方式实现全局最优。本文首先分析了该算法在函数空间中的收敛情况,并证明了谱图理论可以以类似于数值向量的方式应用于函数空间。在此基础上,提出了基于共识的多跳联合精馏(CMFD)的神经网络(NN)实现元算法。CMFD利用知识蒸馏来实现相邻设备之间的功能聚合,而无需参数平均。CMFD的一个优点是,它甚至可以在分布式学习者之间使用不同的神经网络模型。虽然CMFD不能很好地反映元算法的行为,但对元算法收敛性质的讨论促进了对CMFD的直观理解,仿真评估表明,对于多个任务,使用CMFD可以收敛神经网络模型。仿真结果还表明,对于弱连接网络,CMFD比参数聚合具有更高的精度,而且CMFD比参数聚合方法更稳定
This paper proposes a fully decentralized federated learning (FL) scheme for Internet of Everything (IoE) devices that are connected via multi-hop networks. Because FL algorithms hardly converge the parameters of machine learning (ML) models, this paper focuses on the convergence of ML models in function spaces. Considering that the representative loss functions of ML tasks e.g., mean squared error (MSE) and Kullback-Leibler (KL) divergence, are convex functionals, algorithms that directly update functions in function spaces could converge to the optimal solution. The key concept of this paper is to tailor a consensus-based optimization algorithm to work in the function space and achieve the global optimum in a distributed manner. This paper first analyzes the convergence of the proposed algorithm in a function space, which is referred to as a meta-algorithm, and shows that the spectral graph theory can be applied to the function space in a manner similar to that of numerical vectors. Then, consensus-based multi-hop federated distillation (CMFD) is developed for a neural network (NN) to implement the meta-algorithm. CMFD leverages knowledge distillation to realize function aggregation among adjacent devices without parameter averaging. An advantage of CMFD is that it works even with different NN models among the distributed learners. Although CMFD does not perfectly reflect the behavior of the meta-algorithm, the discussion of the meta-algorithm's convergence property promotes an intuitive understanding of CMFD, and simulation evaluations show that NN models converge using CMFD for several tasks. The simulation results also show that CMFD achieves higher accuracy than parameter aggregation for weakly connected networks, and CMFD is more stable than parameter aggregation methods