Communication Efficient Parallel Reinforcement Learning

Communication Efficient Parallel Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Mridul Agarwal;Bhargav Ganguly;V. Aggarwal
Mridul Agarwal;Bhargav Ganguly;V. Aggarwal
中科院分区:
其他
文献类型:
--
作者:
Mridul Agarwal;Bhargav Ganguly;V. Aggarwal

文献摘要

被引文献

相似文献

我们考虑的问题,M代理与M相同的和独立的环境与S状态和A行动使用强化学习T轮。代理与中央服务器共享数据,以尽量减少他们的遗憾。我们的目标是找到一种算法,允许代理人在不频繁的通信回合中最大限度地减少遗憾。我们提供了在每个代理处运行的DIST-UCRL,并证明了M个代理的总累积遗憾的上界为DSTO(DS_MAT),对于直径为D,状态数为S,动作数为A的马尔可夫决策过程。代理同步后,他们的访问任何状态动作对超过一定的阈值。利用这一点,我们得到了一个绑定的O(MSA log(MT))的总数量的通信轮。最后,我们针对多种环境评估了该算法,并证明了该算法与UCRL 2算法的始终通信版本相当,而通信量明显较低。
We consider the problem where M agents interact with M identical and independent environments with S states and A actions using reinforcement learning for T rounds. The agents share their data with a central server to minimize their regret. We aim to find an algorithm that allows the agents to minimize the regret with infrequent communication rounds. We provide DIST -UCRL which runs at each agent and prove that the total cumulative regret of M agents is upper bounded as ˜ O ( DS √ MAT ) for a Markov Decision Process with diameter D , number of states S , and number of actions A . The agents synchronize after their visitations to any state-action pair exceeds a cer-tain threshold. Using this, we obtain a bound of O ( MSA log( MT )) on the total number of communications rounds. Finally, we evaluate the algorithm against multiple environments and demonstrate that the proposed algorithm performs at par with an always communication version of the UCRL2 algorithm, while with significantly lower communication.