InitLight: Initial Model Generation for Traffic Signal Control Using Adversarial Inverse Reinforcement Learning

InitLight: Initial Model Generation for Traffic Signal Control Using Adversarial Inverse Reinforcement Learning
复制标题

DOI:
10.24963/ijcai.2023/550
复制
发表时间:
2023-08
期刊:
--
影响因子:
--
通讯作者:
Yutong Ye;Yingbo Zhou;Jiepin Ding;Ting Wang;Mingsong Chen;Xiang Lian
Yutong Ye;Yingbo Zhou;Jiepin Ding;Ting Wang;Mingsong Chen;Xiang Lian
中科院分区:
其他
文献类型:
--
作者:
Yutong Ye;Yingbo Zhou;Jiepin Ding;Ting Wang;Mingsong Chen;Xiang Lian

文献摘要

相似文献

现有的基于强化学习(RL)的交通信号控制(TSC)方法在策略学习过程中,由于智能体与固定交通环境之间的反复试错式交互,存在RL训练时间长、RL智能体对其他复杂交通环境适应性差等问题。针对这些问题,我们提出了一种新的基于对抗性反向强化学习(AIR)的预训练方法InitLight,该方法能够有效地为TSC代理生成初始模型。与传统的基于RL的TSC方法针对特定的多交叉环境同时训练大量代理不同,InitLight仅基于多个单交叉环境及其专家轨迹预先训练一个单一初始模型。由于InitLight学习的奖励函数可以最优地恢复不同交叉口的地面真实TSC奖励,因此预先训练的代理可以作为初始模型部署在任何交通环境的交叉口,以加速后续的全局RL训练。综合实验结果表明,由InitLight生成的初始模型不仅能够以较少的场景数显著加快收敛速度,而且具有较强的泛化能力,能够适应各种复杂的交通环境。
Due to repetitive trial-and-error style interactions between agents and a fixed traffic environment during the policy learning, existing Reinforcement Learning (RL)-based Traffic Signal Control (TSC) methods greatly suffer from long RL training time and poor adaptability of RL agents to other complex traffic environments. To address these problems, we propose a novel Adversarial Inverse Reinforcement Learning (AIRL)-based pre-training method named InitLight, which enables effective initial model generation for TSC agents. Unlike traditional RL-based TSC approaches that train a large number of agents simultaneously for a specific multi-intersection environment, InitLight pre-trains only one single initial model based on multiple single-intersection environments together with their expert trajectories. Since the reward function learned by InitLight can recover ground-truth TSC rewards for different intersections at optimality, the pre-trained agent can be deployed at intersections of any traffic environments as initial models to accelerate subsequent overall global RL training. Comprehensive experimental results show that, the initial model generated by InitLight can not only significantly accelerate the convergence with much fewer episodes, but also own superior generalization ability to accommodate various kinds of complex traffic environments.