Federated Multi-armed Bandits with Personalization

Federated Multi-armed Bandits with Personalization
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Chengshuai Shi;Cong Shen;Jing Yang
Chengshuai Shi;Cong Shen;Jing Yang
中科院分区:
其他
文献类型:
--
作者:
Chengshuai Shi;Cong Shen;Jing Yang

文献摘要

相似文献

提出了一种通用的个性化联邦多臂强盗(PF-MAB)框架,它是一种类似于监督学习中的联邦学习(FL)框架的新的强盗范式,具有FL的个性化特征。在PF-MAB框架下,研究了一种灵活平衡泛化和个性化的混合强盗学习问题。给出了混合模型的下界分析。然后,我们提出了个性化的联邦置信上限(PF-UCB)算法,其中的探索长度是精心选择,以实现所需的平衡学习的局部模型和提供全局信息的混合学习目标。理论分析证明,PF-UCB算法在不考虑个性化程度的情况下,都能达到O(\log(T))$的遗憾度,并且具有与下界相似的实例依赖性.实验结果验证了理论分析的正确性,并证明了算法的有效性。
A general framework of personalized federated multi-armed bandits (PF-MAB) is proposed, which is a new bandit paradigm analogous to the federated learning (FL) framework in supervised learning and enjoys the features of FL with personalization. Under the PF-MAB framework, a mixed bandit learning problem that flexibly balances generalization and personalization is studied. A lower bound analysis for the mixed model is presented. We then propose the Personalized Federated Upper Confidence Bound (PF-UCB) algorithm, where the exploration length is chosen carefully to achieve the desired balance of learning the local model and supplying global information for the mixed learning objective. Theoretical analysis proves that PF-UCB achieves an $O(\log(T))$ regret regardless of the degree of personalization, and has a similar instance dependency as the lower bound. Experiments using both synthetic and real-world datasets corroborate the theoretical analysis and demonstrate the effectiveness of the proposed algorithm.