Reinforcement Learning with Probabilistic Guarantees for Autonomous Driving

Reinforcement Learning with Probabilistic Guarantees for Autonomous Driving
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Maxime Bouton;J. Karlsson;A. Nakhaei;K. Fujimura;Mykel J. Kochenderfer;Jana Tumova
Maxime Bouton;J. Karlsson;A. Nakhaei;K. Fujimura;Mykel J. Kochenderfer;Jana Tumova
中科院分区:
其他
文献类型:
--
作者:
Maxime Bouton;J. Karlsson;A. Nakhaei;K. Fujimura;Mykel J. Kochenderfer;Jana Tumova

文献摘要

被引文献

相似文献

为城市自动驾驶设计可靠的决策策略具有挑战性。强化学习(RL)已被用于在不确定的环境中自动导出合适的行为,但它不能保证所产生的策略的性能。我们提出了一个通用的方法来执行概率保证的RL代理。一个探索策略是在训练之前推导出来的,该策略约束代理在满足用线性时序逻辑(LTL)表示的期望概率规范的动作之间进行选择。将搜索空间缩小到满足LTL公式的策略有助于训练并简化奖励设计。本文概述了一个交叉口的情况下,涉及多个交通参与者的案例研究。由此产生的政策优于基于规则的启发式方法的效率,同时表现出强有力的安全保证。
Designing reliable decision strategies for autonomous urban driving is challenging. Reinforcement learning (RL) has been used to automatically derive suitable behavior in uncertain environments, but it does not provide any guarantee on the performance of the resulting policy. We propose a generic approach to enforce probabilistic guarantees on an RL agent. An exploration strategy is derived prior to training that constrains the agent to choose among actions that satisfy a desired probabilistic specification expressed with linear temporal logic (LTL). Reducing the search space to policies satisfying the LTL formula helps training and simplifies reward design. This paper outlines a case study of an intersection scenario involving multiple traffic participants. The resulting policy outperforms a rule-based heuristic approach in terms of efficiency while exhibiting strong guarantees on safety.