Parallel reward and punishment control in humans and robots: Safe reinforcement learning using the MaxPain algorithm
Parallel reward and punishment control in humans and robots: Safe reinforcement learning using the MaxPain algorithm
复制标题
人类和机器人的并行奖励和惩罚控制:使用 MaxPain 算法的安全强化学习
DOI:
10.1109/devlrn.2017.8329799
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
B. Seymour
中科院分区:
文献类型:
--
作者:
Stefan Elfwing;B. Seymour
An important issue in reinforcement learning systems for autonomous agents is whether it makes sense to have separate systems for predicting rewards and punishments. In robotics, learning and control are typically achieved by a single controller, with punishments coded as negative rewards. However in biological systems, some evidence suggests that the brain has a separate system for punishment. Although this may in part be due to biological constraints of implementing negative quantities, it raises the question as to whether there is any computational rationale for keeping reward and punishment prediction operationally distinct. Here we outline a basic argument supporting this idea, based on the proposition that learning best-case predictions (as in Q-learning) does not always achieve the safest behaviour. We introduce a modified RL scheme involving a new algorithm which we call ’MaxPain’ — which back-ups worst-case predictions in parallel, and then scales the two predictions in a multiattribute RL policy. i.e. independently learning ‘what to do’ as well as ‘what not to do’ and then combining this information. We show how this scheme can improve performance in benchmark RL environments, including a grid-world experiment and delayed version of the mountain car experiment. In particular, we demonstrate how early exploration and learning are substantially improved, leading to much ‘safer’ behaviour. In conclusion, the results illustrate the importance of independent punishment prediction in RL, and provide a testable framework for better understanding punishment (such as pain) and avoidance in humans, in both health and disease.