Multi-player multi-armed bandits: Decentralized learning with IID rewards
Multi-player multi-armed bandits: Decentralized learning with IID rewards
复制标题
多人多臂强盗:具有 IID 奖励的去中心化学习
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
R. Jain
中科院分区:
文献类型:
--
作者:
D. Kalathil;Naumaan Nayyar;R. Jain
We consider the decentralized multi-armed bandit problem with distinct arms for each players. Each player can pick one arm at each time instant and can get a random reward from an unknown distribution with an unknown mean. The arms give different rewards to different players. If more than one player select the same arm, everyone gets a zero reward. There is no dedicated control channel for communication or coordination among the user. We propose an online learning algorithm called dUCB4 which achieves a near-O(log2 T). The motivation comes from opportunistic spectrum access by multiple secondary users in cognitive radio networks wherein they must pick among various wireless channels that look different to different users.