Multi-agent Robust Time Differential Reinforcement Learning Over Communicated Networks

Multi-agent Robust Time Differential Reinforcement Learning Over Communicated Networks
复制标题

通信网络上的多智能体鲁棒时间差分强化学习

DOI:
10.23919/chicc.2018.8483961
复制
发表时间:
2018-07
期刊:
2018 37th Chinese Control Conference (CCC)
影响因子:
--
通讯作者:
Xiangmin Han
Xiangmin Han
中科院分区:
其他
文献类型:
--
作者:
Jiahong Li;Nan Ma;Xiangmin Han

文献摘要

参考文献

相似文献

近年来,多智能体强化学习(MARL)的研究在自动驾驶等领域引起了广泛的关注。MARL的主要问题是如何处理环境中的不确定性和连接的智能体之间的交互。针对这一问题,提出了一种分布式鲁棒时间差分深度Q网络算法(MARTD-DQN)。MARTD-DQN由分散MARL算法(DMARL)和鲁棒TD深度Q网络算法(RTD-DQN)两部分组成。DMARL通过融合通信网络上邻居的状态来提高策略估计的鲁棒性。RTD-DQN通过在线估计不确定度,提高了对离群值的鲁棒性。通过将这两种算法相结合,该算法不仅对节点失效具有鲁棒性,而且对离群点也具有鲁棒性。将该算法应用于自动汽车ACC仿真。仿真结果表明了该算法的有效性。
Recently, the researches on multi-agent reinforcement learning (MARL) have attracted tremendous interest in many applications, especially for autonomous driving. The main problem of MARL is how to deal with the uncertainty in the environment and the interaction between the connected agents. To solve the problem, a distributed robust temporal differential deep Q-network algorithm (MARTD-DQN) was developed in this paper. MARTD-DQN consists of two parts, the decentralized MARL algorithm (DMARL) and the robust TD deep Q-network algorithm (RTD-DQN). DMARL improves the robustness of the policy estimation by fusing the states from the neighbors over communicated networks. RTD- DQN improves the robustness to outliers through on-line estimation of the uncertainty. By combining the two algorithms, the proposed algorithm can be robust not only to node failures but also to the outliers. Then the proposed algorithm is applied to ACC simulations of autonomous cars. The simulation results are given to show the efficiency of the proposed algorithm.
DOI: --
发表时间: 1999
期刊: --
影响因子: --
作者:
Vijay R. Konda;J. Tsitsiklis
通讯作者: Vijay R. Konda;J. Tsitsiklis
DOI: 10.2307/2291177
发表时间: 1994-04
影响因子: 4.4
作者:
M. Puterman
通讯作者: M. Puterman
DOI: 10.1109/tcyb.2017.2655725
发表时间: 2017-01
影响因子: 11.8
作者:
H. Choi;C. Ahn;H. Karimi;M. Lim
通讯作者: H. Choi;C. Ahn;H. Karimi;M. Lim
DOI: 10.1287/mnsc.1060.0614
发表时间: 2007-02
期刊: Manag. Sci.
影响因子: --
作者:
Shie Mannor;D. Simester;Peng Sun;J. Tsitsiklis
通讯作者: Shie Mannor;D. Simester;Peng Sun;J. Tsitsiklis
DOI: 10.1177/105345129302800510
发表时间: 1993-05
影响因子: 0.8
作者:
Judith Hylton
通讯作者: Judith Hylton