Faithful and Effective Reward Schemes for Model-Free Reinforcement Learning of Omega-Regular Objectives

Faithful and Effective Reward Schemes for Model-Free Reinforcement Learning of Omega-Regular Objectives
复制标题

DOI:
10.1007/978-3-030-59152-6_6
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
E. M. Hahn;Mateo Perez;S. Schewe;F. Somenzi;Ashutosh Trivedi;D. Wojtczak
E. M. Hahn;Mateo Perez;S. Schewe;F. Somenzi;Ashutosh Trivedi;D. Wojtczak
中科院分区:
其他
文献类型:
--
作者:
E. M. Hahn;Mateo Perez;S. Schewe;F. Somenzi;Ashutosh Trivedi;D. Wojtczak

文献摘要

相似文献

使用线性时间时序逻辑或各种形式的欧米茄自动机指定的欧米茄正则属性在指定强化学习(RL)的目标中越来越多地使用。出现的关键问题是将目标忠实有效地转化为无模型RL的标量奖励。最近的一种方法利用Büchi自动机与限制非确定性减少搜索的最优策略的一个正则属性,一个简单的可达性目标。这种转换的一个可能的缺点是可达性奖励是稀疏的,只在每集结束时获得。另一种方法将对最优策略的搜索简化为具有两个相互依赖的折扣参数的优化问题。虽然这种方法提供了比减少可达性更密集的奖励,但它不容易映射到现成的RL算法。我们提出了一个奖励计划,减少了一个单一的折扣参数,产生密集的奖励,并与现成的RL算法的优化问题的最佳策略的搜索。最后,我们报告了一个实验比较这些和其他奖励计划的无模型强化学习与欧米茄定期目标。
Omega-regular properties—specified using linear time temporal logic or various forms of omega-automata—find increasing use in specifying the objectives of reinforcement learning (RL). The key problem that arises is that of faithful and effective translation of the objective into a scalar reward for model-free RL. A recent approach exploits Büchi automata with restricted nondeterminism to reduce the search for an optimal policy for an-regular property to that for a simple reachability objective. A possible drawback of this translation is that reachability rewards are sparse, being reaped only at the end of each episode. Another approach reduces the search for an optimal policy to an optimization problem with two interdependent discount parameters. While this approach provides denser rewards than the reduction to reachability, it is not easily mapped to off-the-shelf RL algorithms. We propose a reward scheme that reduces the search for an optimal policy to an optimization problem with a single discount parameter that produces dense rewards and is compatible with off-the-shelf RL algorithms. Finally, we report an experimental comparison of these and other reward schemes for model-free RL with omega-regular objectives.