Hierarchical Advantage for Reinforcement Learning in Parameterized Action Space
Hierarchical Advantage for Reinforcement Learning in Parameterized Action Space
复制标题
DOI:
10.1109/cog52621.2021.9619068
复制
发表时间:
2021-08
期刊:
影响因子:
--
通讯作者:
Zhejie Hu;Tomoyuki Kaneko
中科院分区:
文献类型:
--
作者:
Zhejie Hu;Tomoyuki Kaneko
We propose a hierarchical architecture for the advantage function to improve the performance of reinforcement learning in parameterized action space, which consists of a set of discrete actions and a set of continuous parameters corresponding to each discrete action. The hierarchical architecture extends the actor-critic architecture with two specialized advantage functions, one for discrete actions and the other for continuous parameters, to estimate a better baseline. We incorporate this architecture into proximal policy optimization, which is referred to as HA-PPO. We evaluated all of our methods on the Half Field Offense domain, and found that the hierarchical architecture of the advantage function, which is referred to as the hierarchical advantage, helps to stabilize the learning and leads to a better performance.