Distributed Multiarmed Bandits

Distributed Multiarmed Bandits
复制标题

分布式多臂强盗

DOI:
10.1109/tac.2023.3247982
复制
发表时间:
2023
影响因子:
6.8
通讯作者:
Liu, Ji
Liu, Ji
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhu, Jingxuan;Liu, Ji

文献摘要

参考文献

相似文献

本文研究了一类具有异质报酬观测值的分布式多臂强盗问题。该问题是合作解决的代理人假设每个代理人面临一个共同的一组ofarms,但只观察到局部有偏见的奖励的武器。每个代理的目标是最小化相对于手臂的真实奖励的累积预期后悔,其中每个手臂的真实奖励的平均值等于所有代理的观察到的偏差奖励的平均值。每个代理递归地更新其决策,利用其邻居的信息。邻居关系由一个时间依赖的有向图描述,其顶点对应于代理,其弧描绘邻居关系。将经典的分布式平均算法与著名的置信上界bandit算法相结合,提出了一种全分布式bandit算法。它表明,对于任何一致强连通序列,该算法实现了保证后悔的顺序为每个代理。
This article studies a distributed multiarmed bandit problem with heterogeneous observations of rewards. The problem is cooperatively solved byagents assuming each agent faces a common set ofarms yet observes only local biased rewards of the arms. The goal of each agent is to minimize the cumulative expected regret with respect to the true rewards of the arms, where the mean of each arm's true reward equals the average of the means of all agents' observed biased rewards. Each agent recursively updates its decision by utilizing the information from its neighbors. Neighbor relationships are described by a time-dependent directed graphwhose vertices correspond to agents and whose arcs depict neighbor relationships. A fully distributed bandit algorithm is proposed, which couples the classical distributed averaging algorithm and the celebrated upper confidence bound bandit algorithm. It is shown that for any uniformly strongly connected sequence of, the algorithm achieves guaranteed regret for each agent at the order of.
分布式多人强盗 - 权力的游戏方法
DOI: --
发表时间: 2018
期刊: Neural Information Processing Systems
影响因子: --
作者:
Ilai Bistritz;Amir Leshem
通讯作者: Amir Leshem
DOI: --
发表时间: 2021-10
期刊: --
影响因子: --
作者:
Ruiquan Huang;Weiqiang Wu;Jing Yang;Cong Shen
通讯作者: Ruiquan Huang;Weiqiang Wu;Jing Yang;Cong Shen
DOI: --
发表时间: 2020-10
期刊: ArXiv
影响因子: --
作者:
Abhimanyu Dubey;A. Pentland
通讯作者: Abhimanyu Dubey;A. Pentland
多代理多武装强盗中的社会学习
DOI: 10.1145/3393691.3394217
发表时间: 2019
期刊: Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子: --
作者:
Sankararaman, Abishek;Ganesh, Ayalvadi;Shakkottai, Sanjay
通讯作者: Shakkottai, Sanjay
DOI: --
发表时间: 2020
期刊: IEEE Conference on Decision and Control
影响因子: --
作者:
Jingxuan Zhu;Romeil Sandhu;Ji Liu
通讯作者: Ji Liu