Policy Optimization with Linear Temporal Logic Constraints
Policy Optimization with Linear Temporal Logic Constraints
复制标题
DOI:
10.48550/arxiv.2206.09546
复制
发表时间:
2022-06
期刊:
影响因子:
--
通讯作者:
Cameron Voloshin;Hoang Minh Le;Swarat Chaudhuri;Yisong Yue
中科院分区:
文献类型:
--
作者:
Cameron Voloshin;Hoang Minh Le;Swarat Chaudhuri;Yisong Yue
We study the problem of policy optimization (PO) with linear temporal logic (LTL) constraints. The language of LTL allows flexible description of tasks that may be unnatural to encode as a scalar cost function. We consider LTL-constrained PO as a systematic framework, decoupling task specification from policy selection, and as an alternative to the standard of cost shaping. With access to a generative model, we develop a model-based approach that enjoys a sample complexity analysis for guaranteeing both task satisfaction and cost optimality (through a reduction to a reachability problem). Empirically, our algorithm can achieve strong performance even in low-sample regimes.