Asynchronous Upper Confidence Bound Algorithms for Federated Linear Bandits

Asynchronous Upper Confidence Bound Algorithms for Federated Linear Bandits
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Chuanhao Li;Hongning Wang
Chuanhao Li;Hongning Wang
中科院分区:
其他
文献类型:
--
作者:
Chuanhao Li;Hongning Wang

文献摘要

被引文献

相似文献

线性上下文强盗是一个流行的在线学习问题。它主要是在集中学习环境中进行研究的。随着大规模去中心化模型学习(例如联邦学习)的需求激增,如何在降低沟通成本的同时保持遗憾最小化成为一个开放的挑战。在本文中,我们研究联邦学习环境中的线性上下文老虎机。我们分别提出了一个具有异步模型更新和通信功能的通用框架,用于同构客户端和异构客户端的集合。对这种分布式学习框架下的遗憾和沟通成本进行了严谨的理论分析;广泛的实证评估证明了我们解决方案的有效性。
Linear contextual bandit is a popular online learning problem. It has been mostly studied in centralized learning settings. With the surging demand of large-scale decentralized model learning, e.g., federated learning, how to retain regret minimization while reducing communication cost becomes an open challenge. In this paper, we study linear contextual bandit in a federated learning setting. We propose a general framework with asynchronous model update and communication for a collection of homogeneous clients and heterogeneous clients, respectively. Rigorous theoretical analysis is provided about the regret and communication cost under this distributed learning framework; and extensive empirical evaluations demonstrate the effectiveness of our solution.