Combined Longitudinal and Lateral Control of Autonomous Vehicles based on Reinforcement Learning

Combined Longitudinal and Lateral Control of Autonomous Vehicles based on Reinforcement Learning
复制标题

基于强化学习的自动驾驶车辆纵向横向联合控制

DOI:
--
复制
发表时间:
2021
期刊:
American Control Conference
影响因子:
--
通讯作者:
Zhong
Zhong
中科院分区:
--
文献类型:
--
作者:
Leilei Cui;Kaan Özbay;Zhong

文献摘要

被引文献

相似文献

本文提出了一种数据驱动的最优控制方法,以使自主车辆与前车保持期望的距离并保持在车道上。首先,推导了自主车辆的动力学。为了克服前沿限制,定义垂直于在前车辆的虚拟在前车辆。跟踪误差被定义为自主车辆的前视点与虚拟前方车辆之间的偏差。然后,推导出误差系统。其次,基于误差系统,以跟踪误差和能量消耗所决定的代价最小为目标,建立了系统的Hamilton-Jacobi-Bellman(HJB)方程。提出了一种基于模型的策略迭代技术来求解HJB方程。第三,提出了一种两阶段数据驱动的策略迭代算法,并利用自适应动态规划(ADP)实现了该算法。计算机仿真验证了所提出的数据驱动的最优控制方法的有效性。
In this paper, in order for the autonomous vehicle to keep a desired distance from the preceding vehicle and stay in the lane, a data-driven optimal control approach is proposed. Firstly, the dynamics of the autonomous vehicle is derived. In order to overcome the cutting-edge limitation, a virtual preceding vehicle is defined which is perpendicular to the preceding vehicle. The tracking error is defined as the deviation between the look ahead point of the autonomous vehicle and the virtual preceding vehicle. Then, the error system is derived. Secondly, based on the error system, in order to minimize the cost determined by the tracking error and the energy consumption, the Hamilton-Jacobi-Bellman (HJB) equation is established. A model-based policy iteration technique is proposed to solve the HJB equation. Thirdly, a two-phase data-driven policy iteration algorithm is proposed and implemented by using adaptive dynamic programming (ADP). The efficacy of the proposed data-driven optimal control approach is validated by computer simulations.