Markov Decision Processes with Arbitrary Reward Processes

Markov Decision Processes with Arbitrary Reward Processes
复制标题

DOI:
10.1287/moor.1090.0397
复制
发表时间:
2008-11
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
Jia Yuan Yu;Shie Mannor;N. Shimkin
Jia Yuan Yu;Shie Mannor;N. Shimkin
中科院分区:
其他
文献类型:
--
作者:
Jia Yuan Yu;Shie Mannor;N. Shimkin

文献摘要

被引文献

相似文献

我们考虑一个学习问题,决策者与一个标准的马尔可夫决策过程,除了奖励函数随时间任意变化。我们表明,对每一个可能实现的奖励过程中,代理可以执行以及-在事后-作为每一个固定的政策。这推广了经典的重复博弈的无悔结果。具体来说,我们提出了一个有效的在线算法-在强化学习的精神-,确保代理的平均性能损失随着时间的推移消失,提供的环境是不经意的代理的行动。此外,还可以修改基本算法,以科普奖励观察仅限于代理轨迹的情况。我们提出了进一步的修改,通过使用函数近似和跟踪最优策略,通过不频繁的变化,降低计算成本。
We consider a learning problem where the decision maker interacts with a standard Markov decision process, with the exception that the reward functions vary arbitrarily over time. We show that, against every possible realization of the reward process, the agent can perform as well---in hindsight---as every stationary policy. This generalizes the classical no-regret result for repeated games. Specifically, we present an efficient online algorithm---in the spirit of reinforcement learning---that ensures that the agent's average performance loss vanishes over time, provided that the environment is oblivious to the agent's actions. Moreover, it is possible to modify the basic algorithm to cope with instances where reward observations are limited to the agent's trajectory. We present further modifications that reduce the computational cost by using function approximation and that track the optimal policy through infrequent changes.