Two Dimensional Evaluation Reinforcement Learning
Two Dimensional Evaluation Reinforcement Learning
复制标题
二维评估强化学习
DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
T. Omori
中科院分区:
文献类型:
--
作者:
Hiroyuki Okada;H. Yamakawa;T. Omori
To solve the problem of tradeoff between exploration and exploitation actions in reinforcement learning, the authors have proposed two-dimensional evaluation reinforcement learning, which distinguishes between reward and punishment evaluation forecasts. The proposed method use these difference between reward evaluation and punishment evaluation as a factor for determining the action and the sum as a parameter for determining the ratio of exploration to exploitation. In this paper we described an experiment with a mobile robot searching for a path and the subsequent conflict between exploration and exploitation actions. The results of the experiment pro ve that using the proposed method of reinforcement learning using the two dimensions of reward and punishment can generate a better path than using the conventional reinforcement learning method.