Translating Omega-Regular Specifications to Average Objectives for Model-Free Reinforcement Learning

Translating Omega-Regular Specifications to Average Objectives for Model-Free Reinforcement Learning
复制标题

DOI:
10.5555/3535850.3535933
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
M. Kazemi;Mateo Perez;F. Somenzi;Sadegh Soudjani;Ashutosh Trivedi;Alvaro Velasquez
M. Kazemi;Mateo Perez;F. Somenzi;Sadegh Soudjani;Ashutosh Trivedi;Alvaro Velasquez
中科院分区:
其他
文献类型:
--
作者:
M. Kazemi;Mateo Perez;F. Somenzi;Sadegh Soudjani;Ashutosh Trivedi;Alvaro Velasquez

文献摘要

相似文献

最近强化学习(RL)的成功重新引起了人们对奖励函数设计的关注,通过奖励函数来加强或阻止代理行为。手动设计奖励函数既繁琐又容易出错。另一种方法是指定一个形式化的、明确的逻辑要求,它会自动转换为一个可供学习的奖励函数。Omega-正则语言是线性时态逻辑(LTL)的一个子集,由于它们在验证和综合中的使用,因此是指定此类需求的自然选择。然而,基于omega-Regular语言的当前技术以插段式方式学习,从而在学习期间将环境周期性地重置到初始状态。在某些情况下,这一假设具有挑战性,甚至不可能得到满足。相反,在持续设置中,代理在单个生命周期内探索环境,而不进行重置。对于在无限的代理行为痕迹上定义的omega规则规范进行推理,这是一个更自然的设置。在这种情况下,优化平均奖励而不是通常的折扣奖励更自然,因为无限范围的目标对折扣RL解的收敛提出了挑战。我们将注意力限制在符合绝对活跃度规范的omega-Regular语言上。根据持续问题的精神,这些规范不能因代理行为的任何有限前缀而无效。我们提出了一个从绝对活跃度omega-正则语言到RL的平均奖励目标的转换。我们的简化可以在不完全了解环境的情况下即时完成,从而能够使用无模型的RL算法。此外,我们提出了一种奖励结构,与以往的方法不同,该结构使得RL在传递MDP时不需要间歇性重置。通过不同的基准测试,我们提出的对omega-Regular规范定义的持续任务使用平均报酬RL的方法比利用折扣RL的竞争方法更有效。
Recent success in reinforcement learning (RL) has brought renewed attention to the design of reward functions by which agent behavior is reinforced or deterred. Manually designing reward functions is tedious and error-prone. An alternative approach is to specify a formal, unambiguous logic requirement, which is automatically translated into a reward function to be learned from. Omega-regular languages, of which Linear Temporal Logic (LTL) is a subset, are a natural choice for specifying such requirements due to their use in verification and synthesis. However, current techniques based on omega-regular languages learn in an episodic manner whereby the environment is periodically reset to an initial state during learning. In some settings, this assumption is challenging or impossible to satisfy. Instead, in the continuing setting the agent explores the environment without resets over a single lifetime. This is a more natural setting for reasoning about omega-regular specifications defined over infinite traces of agent behavior. Optimizing the average reward instead of the usual discounted reward is more natural in this case due to the infinite-horizon objective that poses challenges to the convergence of discounted RL solutions. We restrict our attention to the omega-regular languages which correspond to absolute liveness specifications. These specifications cannot be invalidated by any finite prefix of agent behavior, in accordance with the spirit of a continuing problem. We propose a translation from absolute liveness omega-regular languages to an average reward objective for RL. Our reduction can be done on-the-fly, without full knowledge of the environment, thereby enabling the use of model-free RL algorithms. Additionally, we propose a reward structure that enables RL without episodic resetting in communicating MDPs, unlike previous approaches. We demonstrate empirically with various benchmarks that our proposed method of using average reward RL for continuing tasks defined by omega-regular specifications is more effective than competing approaches that leverage discounted RL.