Dense reinforcement learning for safety validation of autonomous vehicles

Dense reinforcement learning for safety validation of autonomous vehicles
复制标题

DOI:
10.1038/s41586-023-05732-2
复制
发表时间:
2023-03-23
期刊:
影响因子:
64.8
通讯作者:
Liu, Henry X.
Liu, Henry X.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Feng, Shuo;Sun, Haowei;Liu, Henry X.

文献摘要

被引文献

相似文献

阻碍自动驾驶汽车开发和部署的一个关键瓶颈是,由于安全关键事件的罕见性,在自然驾驶环境中验证其安全性所需的经济和时间成本非常高。在这里,我们报告了一个智能测试环境的开发,在这个环境中,基于人工智能的后台代理被训练来验证自动驾驶汽车在加速模式下的安全性能,而不会失去无偏性。从自然驾驶数据中,背景智能体通过密集的深度强化学习(D2 RL)方法学习执行什么对抗性策略,其中通过删除非安全关键状态并重新连接关键状态来编辑马尔可夫决策过程,以便信息训练数据被加密。D2RL使神经网络能够从具有安全关键事件的密集信息中学习,并实现传统深度强化学习方法难以处理的任务。我们证明了我们的方法的有效性,通过测试一个高度自动化的车辆在高速公路和城市测试轨道与增强现实环境,结合模拟背景车辆与物理道路基础设施和真实的自主测试车辆。我们的研究结果表明,经过D2RL训练的智能体可以将评估过程加速多个数量级(快103到105倍)。此外,D2RL还将与其他安全关键型自主系统一起进行加速测试和培训。
One critical bottleneck that impedes the development and deployment of autonomous vehicles is the prohibitively high economic and time costs required to validate their safety in a naturalistic driving environment, owing to the rarity of safety-critical events(1). Here we report the development of an intelligent testing environment, where artificial-intelligencebased background agents are trained to validate the safety performances of autonomous vehicles in an accelerated mode, without loss of unbiasedness. From naturalistic driving data, the background agents learn what adversarial manoeuvre to execute through a dense deep-reinforcement-learning (D2RL) approach, in which Markov decision processes are edited by removing non-safety-critical states and reconnecting critical ones so that the information in the training data is densified. D2RL enables neural networks to learn from densified information with safety-critical events and achieves tasks that are intractable for traditional deep-reinforcement-learning approaches. We demonstrate the effectiveness of our approach by testing a highly automated vehicle in both highway and urban test tracks with an augmented-reality environment, combining simulated background vehicles with physical road infrastructure and a real autonomous test vehicle. Our results show that the D2RL-trained agents can accelerate the evaluation process by multiple orders of magnitude (103 to 105 times faster). In addition, D2RL will enable accelerated testing and training with other safety-critical autonomous systems.