Two Dimensional Evaluation Reinforcement Learning

Two Dimensional Evaluation Reinforcement Learning
复制标题

二维评估强化学习

DOI:
--
复制
发表时间:
2001
期刊:
International Work-Conference on Artificial and Natural Neural Networks
影响因子:
--
通讯作者:
T. Omori
T. Omori
中科院分区:
--
文献类型:
--
作者:
Hiroyuki Okada;H. Yamakawa;T. Omori

文献摘要

被引文献

相似文献

为了解决强化学习中探索和利用行为之间的权衡问题,提出了区分奖励和惩罚评价预测的二维评价强化学习。该方法将奖励评价与惩罚评价的差异作为决定行动的因素,将其总和作为决定勘探开发比例的参数。在本文中,我们描述了一个实验与移动的机器人搜索路径和随后的冲突之间的探索和开发行动。实验结果表明,采用基于奖惩两个维度的强化学习方法比传统的强化学习方法能够产生更好的路径。
To solve the problem of tradeoff between exploration and exploitation actions in reinforcement learning, the authors have proposed two-dimensional evaluation reinforcement learning, which distinguishes between reward and punishment evaluation forecasts. The proposed method use these difference between reward evaluation and punishment evaluation as a factor for determining the action and the sum as a parameter for determining the ratio of exploration to exploitation. In this paper we described an experiment with a mobile robot searching for a path and the subsequent conflict between exploration and exploitation actions. The results of the experiment pro ve that using the proposed method of reinforcement learning using the two dimensions of reward and punishment can generate a better path than using the conventional reinforcement learning method.