P4FL: An Architecture for Federating Learning With In-Network Processing

P4FL: An Architecture for Federating Learning With In-Network Processing
复制标题

DOI:
10.1109/access.2023.3318109
复制
发表时间:
2023
期刊:
影响因子:
3.9
通讯作者:
Alessio Sacco;Antonino Angi;G. Marchetto;Flavio Esposito
Alessio Sacco;Antonino Angi;G. Marchetto;Flavio Esposito
中科院分区:
计算机科学3区
文献类型:
--
作者:
Alessio Sacco;Antonino Angi;G. Marchetto;Flavio Esposito

文献摘要

被引文献

相似文献

人工智能 (AI) 和机器学习 (ML) 技术的不断发展也伴随着与训练数据相关的隐私问题。联邦学习(FL)是一种相对较新的部分解决此类问题的方法,该技术仅传输经过训练的神经网络模型的参数而不是数据。尽管 FL 可能带来好处,但这种方法可能会导致同步问题(尤其是在众多物联网设备中应用时),网络和服务器可能会变成瓶颈,并且某些节点的负载可能变得不可持续。为了解决这个问题并减少网络流量,在本文中,我们提出了 P4FL,一种新颖的 FL 架构,它使用网络可编程性范式对 P4 交换机进行编程以计算中间聚合。特别是,我们定义了一个基于 MPLS 的自定义带内协议来承载模型参数,并调整 P4 开关行为来聚合模型梯度。然后我们在Mininet中对P4FL进行了评估,并验证了使用网络节点进行网内模型缓存和梯度聚合有两个优点:第一,它减轻了中央FL服务器的瓶颈效应;二是进一步加快了整个训练进度。
The unceasing development of Artificial Intelligence (AI) and Machine Learning (ML) techniques is growing with privacy problems related to the training data. A relatively recent approach to partially cope with such concerns is Federated Learning (FL), a technique in which only the parameters of the trained neural network models are transferred rather than data. Despite the benefits that FL may provide, such an approach can lead to synchronization issues (especially when applied in the context of numerous IoT devices), the network and the server may turn into bottlenecks, and the load may become unsustainable for some nodes. To solve this issue and reduce the traffic on the network, in this paper, we propose P4FL, a novel FL architecture that uses the paradigm of network programmability to program P4 switches to compute intermediate aggregations. In particular, we defined a custom in-band protocol based on MPLS to carry the model parameters and adapted the P4 switch behavior to aggregate model gradients. We then evaluated P4FL in Mininet and verified that using network nodes for in-network model caching and gradient aggregating has two advantages: first, it alleviates the bottleneck effect of the central FL server; second, it further accelerates the entire training progress.