Powering Multi-Task Federated Learning with Competitive GPU Resource Sharing

Powering Multi-Task Federated Learning with Competitive GPU Resource Sharing
复制标题

DOI:
10.1145/3487553.3524859
复制
发表时间:
2022-04
期刊:
Companion Proceedings of the Web Conference 2022
影响因子:
--
通讯作者:
Yongbo Yu;Fuxun Yu;Zirui Xu
Yongbo Yu;Fuxun Yu;Zirui Xu
中科院分区:
其他
文献类型:
--
作者:
Yongbo Yu;Fuxun Yu;Zirui Xu

文献摘要

相似文献

随着认知应用复杂性的增加,联邦学习涉及到复杂的学习任务。例如,自动驾驶系统同时承担多项任务(例如,检测,分类等),并期望FL保持终身智能参与。然而,我们的分析表明,当在GPU上为多个训练任务部署复合FL模型时,会出现一些问题:(1)由于不同任务的数据分布和相应的模型的倾斜导致学习负载高度不平衡,目前的GPU调度方法缺乏有效的资源分配;(2)因此,现有的FL方案只关注异构数据分布,而不关注运行时计算,无法实际实现最优同步联邦。为了解决这些问题,我们提出了一个全栈FL优化方案,以解决设备内GPU调度和设备间FL协调的多任务训练。具体而言,我们的工作阐述了该研究领域的两个关键见解:(1)竞争性资源共享有利于并行模型的执行,并且提出的“虚拟资源”概念可以有效地表征和指导实际的每任务资源利用和分配。(2)考虑建筑层面的协调,可进一步提高FL。我们的实验表明,FL吞吐量可以显著提升。
Federated learning (FL) nowadays involves compound learning tasks as cognitive applications’ complexity increases. For example, a self-driving system hosts multiple tasks simultaneously (e.g., detection, classification, etc.) and expects FL to retain life-long intelligence involvement. However, our analysis demonstrates that, when deploying compound FL models for multiple training tasks on a GPU, certain issues arise: (1) As different tasks’ skewed data distributions and corresponding models cause highly imbalanced learning workloads, current GPU scheduling methods lack effective resource allocations; (2) Therefore, existing FL schemes, only focusing on heterogeneous data distribution but runtime computing, cannot practically achieve optimally synchronized federation. To address these issues, we propose a full-stack FL optimization scheme to address both intra-device GPU scheduling and inter-device FL coordination for multi-task training. Specifically, our works illustrate two key insights in this research domain: (1) Competitive resource sharing is beneficial for parallel model executions, and the proposed concept of “virtual resource” could effectively characterize and guide the practical per-task resource utilization and allocation. (2) FL could be further improved by taking architectural level coordination into consideration. Our experiments demonstrate that the FL throughput could be significantly escalated.