Tractable Reinforcement Learning of Signal Temporal Logic Objectives

Tractable Reinforcement Learning of Signal Temporal Logic Objectives
复制标题

信号时序逻辑目标的易于处理的强化学习

DOI:
--
复制
发表时间:
2020
期刊:
Conference on Learning for Dynamics & Control
影响因子:
--
通讯作者:
P. Seiler
P. Seiler
中科院分区:
--
文献类型:
--
作者:
Harish K. Venkataraman;Derya Aksaray;P. Seiler

文献摘要

被引文献

相似文献

信号时态逻辑(STL)是一种可表达的语言,用于指定有时间限制的真实机器人任务和安全规范。最近,人们有兴趣通过强化学习(RL)来学习满足STL规范的最优策略。学习满足STL规范通常需要足够长的州历史来计算奖励和下一步行动。对历史的需要导致学习问题的状态空间呈指数增长。因此,对于大多数现实世界的应用程序来说,学习问题变得难以计算。在本文中,我们提出了一种紧凑的方法来捕获新的增广状态空间表示中的状态历史。在新的增广状态空间中,提出了一种近似目标(最大化满意概率),并进行了求解。通过仿真,给出了近似解的性能界,并与已有技术的解进行了比较。
Signal temporal logic (STL) is an expressive language to specify time-bound real-world robotic tasks and safety specifications. Recently, there has been an interest in learning optimal policies to satisfy STL specifications via reinforcement learning (RL). Learning to satisfy STL specifications often needs a sufficient length of state history to compute reward and the next action. The need for history results in exponential state-space growth for the learning problem. Thus the learning problem becomes computationally intractable for most real-world applications. In this paper, we propose a compact means to capture state history in a new augmented state-space representation. An approximation to the objective (maximizing probability of satisfaction) is proposed and solved for in the new augmented state-space. We show the performance bound of the approximate solution and compare it with the solution of an existing technique via simulations.