Reinforcement learning for penalty avoiding policy making

Reinforcement learning for penalty avoiding policy making
复制标题

强化学习避免惩罚政策制定

DOI:
10.1109/icsmc.2000.884990
复制
发表时间:
2000
期刊:
Smc 2000 conference proceedings. 2000 ieee international conference on systems, man and cybernetics. 'cybernetics evolving to systems, humans, organizations, and their complex interactions' (cat. no.0
影响因子:
--
通讯作者:
S. Kobayashi
S. Kobayashi
中科院分区:
--
文献类型:
--
作者:
K. Miyazaki;S. Kobayashi

文献摘要

被引文献

相似文献

强化学习是机器学习的一种。它的目的是使代理适应给定的环境,并提供奖励线索。一般来说,强化学习系统的目的是获得一个最优策略,可以最大化每个动作的预期奖励。但是,它对任何环境都不重要。特别是,如果我们将强化学习应用到工程中,我们希望智能体避免所有惩罚。在马尔可夫决策过程中,我们称规则为惩罚,当且仅当它有惩罚,或者它可以过渡到惩罚状态,在那里它不会得到任何奖励。在抑制所有惩罚规则后,我们的目标是制定一个理性的策略,其每次行动的期望回报大于零。我们提出了避免惩罚的理性决策算法,该算法可以尽可能稳定地抑制任何惩罚,并不断获得奖励。通过将该算法应用于tick-tack-toe问题,证明了该算法的有效性。
Reinforcement learning is a kind of machine learning. It aims to adapt an agent to a given environment with a clue to a reward. In general, the purpose of a reinforcement learning system is to acquire an optimum policy that can maximize expected reward per action. However, it is not always important for any environment. Especially, if we apply reinforcement learning to engineering, we expect the agent to avoid all penalties. In Markov decision processes, we call a rule penalty if and only if it has a penalty or it can transit to a penalty state where it does not contribute to get any reward. After suppressing all penalty rules, we aim to make a rational policy whose expected reward per action is larger than zero. We propose the penalty avoiding rational policy making algorithm that can suppress any penalty as stable as possible and get a reward constantly. By applying the algorithm to the tick-tack-toe its effectiveness is shown.