Hierarchical Advantage for Reinforcement Learning in Parameterized Action Space

Hierarchical Advantage for Reinforcement Learning in Parameterized Action Space
复制标题

DOI:
10.1109/cog52621.2021.9619068
复制
发表时间:
2021-08
期刊:
2021 IEEE Conference on Games (CoG)
影响因子:
--
通讯作者:
Zhejie Hu;Tomoyuki Kaneko
Zhejie Hu;Tomoyuki Kaneko
中科院分区:
其他
文献类型:
--
作者:
Zhejie Hu;Tomoyuki Kaneko

文献摘要

相似文献

我们提出了一个层次结构的优势函数,以提高强化学习的性能参数化的动作空间,其中包括一组离散的动作和一组连续的参数对应于每个离散的动作。分层架构扩展了行动者-批评者架构,具有两个专门的优势函数,一个用于离散动作,另一个用于连续参数,以估计更好的基线。我们将此架构纳入近端策略优化,这是被称为HA-PPO。我们在半场进攻域上评估了我们的所有方法,发现优势函数的分层结构(称为分层优势)有助于稳定学习并获得更好的性能。
We propose a hierarchical architecture for the advantage function to improve the performance of reinforcement learning in parameterized action space, which consists of a set of discrete actions and a set of continuous parameters corresponding to each discrete action. The hierarchical architecture extends the actor-critic architecture with two specialized advantage functions, one for discrete actions and the other for continuous parameters, to estimate a better baseline. We incorporate this architecture into proximal policy optimization, which is referred to as HA-PPO. We evaluated all of our methods on the Half Field Offense domain, and found that the hierarchical architecture of the advantage function, which is referred to as the hierarchical advantage, helps to stabilize the learning and leads to a better performance.