Policy Optimization with Linear Temporal Logic Constraints

Policy Optimization with Linear Temporal Logic Constraints
复制标题

DOI:
10.48550/arxiv.2206.09546
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Cameron Voloshin;Hoang Minh Le;Swarat Chaudhuri;Yisong Yue
Cameron Voloshin;Hoang Minh Le;Swarat Chaudhuri;Yisong Yue
中科院分区:
其他
文献类型:
--
作者:
Cameron Voloshin;Hoang Minh Le;Swarat Chaudhuri;Yisong Yue

文献摘要

相似文献

研究了具有线性时态逻辑(LTL)约束的策略优化问题。LTL语言允许对可能不自然地编码为标量成本函数的任务进行灵活描述。我们将LTL约束的PO作为一个系统框架,将任务说明从策略选择中分离出来,并将其作为成本形成标准的替代。通过使用生成性模型,我们开发了一种基于模型的方法,该方法享受样本复杂性分析,以确保任务满意度和成本最优(通过归结为可达性问题)。实验表明,即使在低样本的情况下,我们的算法也能获得很好的性能。
We study the problem of policy optimization (PO) with linear temporal logic (LTL) constraints. The language of LTL allows flexible description of tasks that may be unnatural to encode as a scalar cost function. We consider LTL-constrained PO as a systematic framework, decoupling task specification from policy selection, and as an alternative to the standard of cost shaping. With access to a generative model, we develop a model-based approach that enjoys a sample complexity analysis for guaranteeing both task satisfaction and cost optimality (through a reduction to a reachability problem). Empirically, our algorithm can achieve strong performance even in low-sample regimes.