Parallel reward and punishment control in humans and robots: Safe reinforcement learning using the MaxPain algorithm

Parallel reward and punishment control in humans and robots: Safe reinforcement learning using the MaxPain algorithm
复制标题

人类和机器人的并行奖励和惩罚控制:使用 MaxPain 算法的安全强化学习

DOI:
10.1109/devlrn.2017.8329799
复制
发表时间:
2017
期刊:
2017 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob)
影响因子:
--
通讯作者:
B. Seymour
B. Seymour
中科院分区:
--
文献类型:
--
作者:
Stefan Elfwing;B. Seymour

文献摘要

被引文献

相似文献

自主代理强化学习系统中的一个重要问题是使用单独的系统来预测奖励和惩罚是否有意义。在机器人技术中,学习和控制通常由单个控制器实现,惩罚被编码为负奖励。然而,在生物系统中,一些证据表明大脑有一个单独的惩罚系统。尽管这可能部分是由于实施负量的生物学限制,但它提出了一个问题:是否存在任何计算原理来保持奖励和惩罚预测在操作上不同。在这里,我们概述了支持这一想法的基本论点,其基础是学习最佳情况预测(如 Q 学习)并不总是能实现最安全的行为。我们引入了一种改进的强化学习方案,其中涉及一种称为“MaxPain”的新算法——它并行支持最坏​​情况的预测,然后在多属性强化学习策略中缩放这两个预测。即独立学习“该做什么”和“不该做什么”,然后结合这些信息。我们展示了该方案如何提高基准 RL 环境中的性能,包括网格世界实验和山地车实验的延迟版本。特别是,我们展示了早期探索和学习如何得到显着改善,从而导致“更安全”的行为。总之,结果说明了强化学习中独立惩罚预测的重要性,并为更好地理解人类健康和疾病方面的惩罚(例如疼痛)和避免提供了可测试的框架。
An important issue in reinforcement learning systems for autonomous agents is whether it makes sense to have separate systems for predicting rewards and punishments. In robotics, learning and control are typically achieved by a single controller, with punishments coded as negative rewards. However in biological systems, some evidence suggests that the brain has a separate system for punishment. Although this may in part be due to biological constraints of implementing negative quantities, it raises the question as to whether there is any computational rationale for keeping reward and punishment prediction operationally distinct. Here we outline a basic argument supporting this idea, based on the proposition that learning best-case predictions (as in Q-learning) does not always achieve the safest behaviour. We introduce a modified RL scheme involving a new algorithm which we call ’MaxPain’ — which back-ups worst-case predictions in parallel, and then scales the two predictions in a multiattribute RL policy. i.e. independently learning ‘what to do’ as well as ‘what not to do’ and then combining this information. We show how this scheme can improve performance in benchmark RL environments, including a grid-world experiment and delayed version of the mountain car experiment. In particular, we demonstrate how early exploration and learning are substantially improved, leading to much ‘safer’ behaviour. In conclusion, the results illustrate the importance of independent punishment prediction in RL, and provide a testable framework for better understanding punishment (such as pain) and avoidance in humans, in both health and disease.