Learning How to Dynamically Route Autonomous Vehicles on Shared Roads

Learning How to Dynamically Route Autonomous Vehicles on Shared Roads
复制标题

DOI:
10.1016/j.trc.2021.103258
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Daniel A. Lazar;Erdem Biyik;Dorsa Sadigh;Ramtin Pedarsani
Daniel A. Lazar;Erdem Biyik;Dorsa Sadigh;Ramtin Pedarsani
中科院分区:
其他
文献类型:
--
作者:
Daniel A. Lazar;Erdem Biyik;Dorsa Sadigh;Ramtin Pedarsani

文献摘要

相似文献

道路拥堵在世界各地造成巨大的成本,而道路网络的干扰,如交通事故,可造成高度拥堵的交通模式。如果一个规划者可以控制网络中所有车辆的路线,他们可以很容易地扭转这种影响。在一个更现实的场景中,我们考虑一个控制自动汽车的规划者,这是所有现有汽车的一小部分。我们研究了一个动态路径博弈,在这个博弈中,自动汽车的路径选择是可控的,而人类驾驶员的反应是自私的和动态的。由于问题太大,我们使用深度强化学习来学习控制自动驾驶汽车的策略。该策略间接影响人类驾驶员以最小化网络拥塞的方式进行路由。为了衡量我们学到的政策的有效性,我们建立了理论结果特征的均衡和经验比较学习的政策结果与最佳可能的均衡。我们证明性质的平衡平行的道路,并提供了一个多项式时间的优化计算最有效的平衡。此外,我们表明,在没有这些政策,高需求和网络扰动将导致大的拥堵,而使用该政策大大减少了旅行时间,最大限度地减少拥堵。据我们所知,这是第一项采用深度强化学习通过间接影响人类在混合自主交通中的路由决策来减少拥堵的工作。
Road congestion induces significant costs across the world, and road network disturbances, such as traffic accidents, can cause highly congested traffic patterns. If a planner had control over the routing of all vehicles in the network, they could easily reverse this effect. In a more realistic scenario, we consider a planner that controls autonomous cars, which are a fraction of all present cars. We study a dynamic routing game, in which the route choices of autonomous cars can be controlled and the human drivers react selfishly and dynamically. As the problem is prohibitively large, we use deep reinforcement learning to learn a policy for controlling the autonomous vehicles. This policy indirectly influences human drivers to route themselves in such a way that minimizes congestion on the network. To gauge the effectiveness of our learned policies, we establish theoretical results characterizing equilibria and empirically compare the learned policy results with best possible equilibria. We prove properties of equilibria on parallel roads and provide a polynomial-time optimization for computing the most efficient equilibrium. Moreover, we show that in the absence of these policies, high demand and network perturbations would result in large congestion, whereas using the policy greatly decreases the travel times by minimizing the congestion. To the best of our knowledge, this is the first work that employs deep reinforcement learning to reduce congestion by indirectly influencing humans’ routing decisions in mixed-autonomy traffic.