Evolution Direction of Reward Appraisal in Reinforcement Learning Agents

Evolution Direction of Reward Appraisal in Reinforcement Learning Agents
复制标题

强化学习智能体奖励评估的演进方向

DOI:
10.1007/978-3-319-92031-3_2
复制
发表时间:
2018
期刊:
Proceedings of the 12th KES International Conference on Agent and Multi-agent Systems: Technologies and Applications
影响因子:
--
通讯作者:
and Nobuhiro Inuzuka
and Nobuhiro Inuzuka
中科院分区:
--
文献类型:
--
作者:
Masaya Miyawaki;Koichi Moriyama;Atsuko Mutoh;Tohgoroh Matsui;and Nobuhiro Inuzuka

文献摘要

相似文献

人类在日常生活中对环境进行评价。我们正在将评估机制引入强化学习代理。我们提出的这种机制之一是基于效用的Q学习,它从代理获得的收益和代理具有的效用导出函数中获得的主观效用来学习行为。在以前的工作中,我们知道,基于支付的进化带来了效用推导功能,促进相互合作的重复囚徒困境游戏。然而,进化过程本身尚未被很好地了解。在这项工作中,我们调查的过程中,是什么决定了演变的方向。我们引入了两个指标,显示基于进化的主观效用的行动偏好,它划分为四个区域的进化空间。在每个地区,指标将解释演变方向。
Humans appraise the environment in daily life. We are implementing appraisal mechanisms into reinforcement learning agents. One of such mechanisms we proposed is the utility-based Q-learning, which learns behaviors from subjective utilities derived from payoffs the agent gains and a utility-derivation function the agent has. In the previous work, we know that payoff-based evolution brings utility-derivation functions that facilitate mutual cooperation in iterated prisoner’s dilemma games. However, the evolution process itself has not yet been known well. In this work, we investigate the process in terms of what determines the evolution direction. We introduce two metrics showing preference of actions based on the evolved subjective utilities, which divide the evolution space into four regions. In each region, the metrics will explain the evolution directions.