Scalable Reinforcement Learning for Multiagent Networked Systems

Scalable Reinforcement Learning for Multiagent Networked Systems
复制标题

DOI:
10.1287/opre.2021.2226
复制
发表时间:
2019-12
期刊:
Oper. Res.
影响因子:
--
通讯作者:
Guannan Qu;A. Wierman;N. Li
Guannan Qu;A. Wierman;N. Li
中科院分区:
其他
文献类型:
--
作者:
Guannan Qu;A. Wierman;N. Li

文献摘要

被引文献

相似文献

以AlphaGo等成功案例为亮点,强化学习(RL)已成为在复杂环境中进行决策的有力工具。然而,到目前为止,强化学习的成功仅限于小规模或单智能体系统。要将强化学习应用于能源、交通和通信网络等大规模网络系统,一个关键的障碍是维度灾难,因为对于这些系统,状态和动作空间可能会随着网络节点数量呈指数级增长。本文试图打破这一维度灾难,并为大型网络系统设计一种可扩展的强化学习方法,名为可扩展演员 - 评论家(SAC)。关键的技术贡献是利用网络结构推导出指数衰减特性,从而能够设计出SAC方法。
Highlighted by success stories like AlphaGo, reinforcement learning (RL) has emerged as a powerful tool for decision making in complex environments. However, the success of RL has thus far been limited to small-scale or single-agent systems. To apply RL to large-scale networked systems such as energy, transportation, and communication networks, a critical hurdle is the curse of dimensionality, because for these systems, the state and action space can be exponentially large in the number of nodes in the network. This article attempts to break this curse of dimensionality and designs a scalable RL method, named scalable actor critic (SAC), for large networked systems. The key technical contribution is to exploit the network structure to derive an exponential decay property, which enables the design of the SAC approach.