FLoX: Federated Learning with FaaS at the Edge

FLoX: Federated Learning with FaaS at the Edge
复制标题

DOI:
10.1109/escience55777.2022.00016
复制
发表时间:
2022-10
期刊:
2022 IEEE 18th International Conference on e-Science (e-Science)
影响因子:
--
通讯作者:
Nikita Kotsehub;Matt Baughman;Ryan Chard;Nathaniel Hudson;Panos Patros;Omer F. Rana;I. Foster;K. Chard
Nikita Kotsehub;Matt Baughman;Ryan Chard;Nathaniel Hudson;Panos Patros;Omer F. Rana;I. Foster;K. Chard
中科院分区:
其他
文献类型:
--
作者:
Nikita Kotsehub;Matt Baughman;Ryan Chard;Nathaniel Hudson;Panos Patros;Omer F. Rana;I. Foster;K. Chard

文献摘要

相似文献

联合学习(FL)是一种分布式机器学习技术,它能够使用孤立的和分布式的数据。利用FL,单独训练单个机器学习模型,然后仅共享和聚集模型参数(例如,神经网络中的权重)以创建全局模型,从而允许数据保持在其原始环境中。虽然许多应用程序可以从FL中受益,但现有的框架并不完整、繁琐且依赖于环境。为了解决这些问题,我们提出了FLoX,这是一个构建在功能X联合无服务器计算平台上的FL框架。FLoX将FL模型培训/推理从基础设施管理中分离出来,从而使用户能够使用一行Python代码在一台或多台远程计算机上轻松部署FL模型。我们使用部署在十个异类和分布式计算终端上的三个基准数据集来评估FLoX。我们表明,FLoX产生的开销很小,特别是在端点之间用于数据传输的大量通信开销方面。我们展示了如何根据参与端点的容量平衡样本和历元的数量,从而在最小程度降低精度的情况下显著减少训练时间。最后,我们表明,全球模型的平均表现比任何单一模型都高出8%。
Federated learning (FL) is a technique for distributed machine learning that enables the use of siloed and distributed data. With FL, individual machine learning models are trained separately and then only model parameters (e.g., weights in a neural network) are shared and aggregated to create a global model, allowing data to remain in its original environment. While many applications can benefit from FL, existing frameworks are incomplete, cumbersome, and environment-dependent. To address these issues, we present FLoX, an FL framework built on the funcX federated serverless computing platform. FLoX decouples FL model training/inference from infrastructure management and thus enables users to easily deploy FL models on one or more remote computers with a single line of Python code. We evaluate FLoX using three benchmark datasets deployed on ten heterogeneous and distributed compute endpoints. We show that FLoX incurs minimal overhead, especially with respect to the large communication overheads between endpoints for data transfer. We show how balancing the number of samples and epochs with respect to the capacities of participating endpoints can significantly reduce training time with minimal reduction in accuracy. Finally, we show that global models consistently outperform any single model on average by 8%.