Model-Free Reinforcement Learning for Branching Markov Decision Processes

Model-Free Reinforcement Learning for Branching Markov Decision Processes
复制标题

用于分支马尔可夫决策过程的无模型强化学习

DOI:
10.1007/978-3-030-81688-9_30
复制
发表时间:
2021
期刊:
Computer Aided Verification. CAV 2021.
影响因子:
--
通讯作者:
Wojtczak, D.
Wojtczak, D.
中科院分区:
--
文献类型:
--
作者:
Hahn, E.M.;Perez, M.;Schewe, S.;Somenzi, F.;Trivedi, A.;Wojtczak, D.

文献摘要

相似文献

我们研究了分支马尔可夫决策过程(BMDPs)的最优控制的强化学习,BMDPs是(多类型)分支马尔可夫链(BMC)的自然扩展。(离散时间)BMC的状态是各种类型的实体的集合,这些实体在产生其他实体的同时产生回报。与BMC相比,其中相同类型的每个实体的演变遵循相同的概率模式,BMDP允许外部控制器从一系列选项中进行选择。这使我们能够研究系统的最佳/最差行为。我们推广了无模型强化学习技术,以计算极限中未知BMDP的最优控制策略。我们提出的实施结果表明,该方法的实用性。
We study reinforcement learning for the optimal control of Branching Markov Decision Processes (BMDPs), a natural extension of (multitype) Branching Markov Chains (BMCs). The state of a (discrete-time) BMCs is a collection of entities of various types that, while spawning other entities, generate a payoff. In comparison with BMCs, where the evolution of a each entity of the same type follows the same probabilistic pattern, BMDPs allow an external controller to pick from a range of options. This permits us to study the best/worst behaviour of the system. We generalise model-free reinforcement learning techniques to compute an optimal control strategy of an unknown BMDP in the limit. We present results of an implementation that demonstrate the practicality of the approach.