Learning to drive from a world on rails

Learning to drive from a world on rails
复制标题

DOI:
10.1109/iccv48922.2021.01530
复制
发表时间:
2021-05
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Di Chen;V. Koltun;Philipp Krähenbühl
Di Chen;V. Koltun;Philipp Krähenbühl
中科院分区:
其他
文献类型:
--
作者:
Di Chen;V. Koltun;Philipp Krähenbühl

文献摘要

相似文献

我们通过基于模型的方法从预先记录的驾驶日志中学习交互式基于视觉的驾驶策略。世界的前向模型监督驾驶策略,预测任何潜在驾驶轨迹的结果。为了支持从预先记录的日志中学习,我们假设世界是在轨道上的,这意味着智能体及其行为都不会影响环境。这个假设大大简化了学习问题,分解成一个非反应的世界模型和一个低维和紧凑的自我车辆的前向模型的动态。我们的方法使用Bellman方程的表格动态编程评估来计算每个训练轨迹的动作值;这些动作值反过来监督最终的基于视觉的驾驶策略。尽管有世界在轨道上的假设,最终驾驶策略在动态和反应性世界中表现良好。它在具有挑战性的CARLA NoCrash基准测试中优于模仿学习以及基于模型和无模型的强化学习。在ProcGen基准测试中,它在导航任务上的样本效率也比最先进的无模型强化学习技术高出一个数量级。
We learn an interactive vision-based driving policy from pre-recorded driving logs via a model-based approach. A forward model of the world supervises a driving policy that predicts the outcome of any potential driving trajectory. To support learning from pre-recorded logs, we assume that the world is on rails, meaning neither the agent nor its actions influence the environment. This assumption greatly simplifies the learning problem, factorizing the dynamics into a non-reactive world model and a low-dimensional and compact forward model of the ego-vehicle. Our approach computes action-values for each training trajectory using a tabular dynamic-programming evaluation of the Bellman equations; these action-values in turn supervise the final vision-based driving policy. Despite the world-on-rails assumption, the final driving policy acts well in a dynamic and reactive world. It outperforms imitation learning as well as model-based and model-free reinforcement learning on the challenging CARLA NoCrash benchmark. It is also an order of magnitude more sample-efficient than state-of-the-art model-free reinforcement learning techniques on navigational tasks in the ProcGen benchmark.