Model-Free Reinforcement Learning for Branching Markov Decision Processes
Model-Free Reinforcement Learning for Branching Markov Decision Processes
复制标题
用于分支马尔可夫决策过程的无模型强化学习
DOI:
10.1007/978-3-030-81688-9_30
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Wojtczak, D.
中科院分区:
文献类型:
--
作者:
Hahn, E.M.;Perez, M.;Schewe, S.;Somenzi, F.;Trivedi, A.;Wojtczak, D.
We study reinforcement learning for the optimal control of Branching Markov Decision Processes (BMDPs), a natural extension of (multitype) Branching Markov Chains (BMCs). The state of a (discrete-time) BMCs is a collection of entities of various types that, while spawning other entities, generate a payoff. In comparison with BMCs, where the evolution of a each entity of the same type follows the same probabilistic pattern, BMDPs allow an external controller to pick from a range of options. This permits us to study the best/worst behaviour of the system. We generalise model-free reinforcement learning techniques to compute an optimal control strategy of an unknown BMDP in the limit. We present results of an implementation that demonstrate the practicality of the approach.