Causal inference multi-agent reinforcement learning for traffic signal control

Causal inference multi-agent reinforcement learning for traffic signal control
复制标题

DOI:
10.1016/j.inffus.2023.02.009
复制
发表时间:
2023-02
期刊:
Inf. Fusion
影响因子:
--
通讯作者:
Shantian Yang;Bo Yang;Zheng Zeng;Zhongfeng Kang
Shantian Yang;Bo Yang;Zheng Zeng;Zhongfeng Kang
中科院分区:
其他
文献类型:
--
作者:
Shantian Yang;Bo Yang;Zheng Zeng;Zhongfeng Kang

文献摘要

相似文献

多智能体强化学习在交通信号控制中的主要挑战是在非平稳多智能体交通环境中产生有效的合作交通信号策略。然而,相邻智能体的时变交通信号策略会使每个智能体受到局部非平稳交通环境的影响;同时,不同的智能体也会产生时变交通信号策略,这进一步导致整个交通环境的非平稳性,因此这些产生的交通信号策略可能是无效的。在这项工作中,我们提出了一个因果推理多智能体强化学习(CI-MA)算法,它可以减轻非平稳的多智能体交通环境的特征表示和优化,最终有助于产生有效的合作交通信号政策。具体而言,首先设计了一个因果推理(CI)模型,通过获取特征表示分布和推导变分下界(即,目标函数);然后,基于所设计的CI模型,提出了一种CI-MA算法,该算法从多Agent交通环境的非平稳性出发,在任务级和时间步级上获取特征表示,并利用获取的特征表示为多Agent产生协作的交通信号策略和Q值;最后,相应的目标函数优化整个算法从因果推理和多智能体强化学习。在不同的非平稳多智能体交通环境中进行实验。实验结果表明,CI-MA算法的性能优于其他现有算法,并证明了该算法可以有效地应用于具有非平稳性的合成和真实业务环境。
A primary challenge in multi-agent reinforcement learning for traffic signal control is to produce effective cooperative traffic-signal policies in non-stationary multi-agent traffic environments. However, each agent suffers from its local non-stationary traffic environment caused by the time-varying traffic-signal policies of adjacent agents; At the same time, different agents also produce time-varying traffic-signal policies, which further results in the non-stationarity of the whole traffic environment, so these produced traffic-signal policies may be ineffective. In this work, we propose a Causal Inference Multi-Agent reinforcement learning (CI-MA) algorithm, which can alleviate the non-stationarity of multi-agent traffic environments from both feature representation and optimization, eventually helps to produce effective cooperative traffic-signal policies. Specifically, a Causal-Inference (CI) model is first designed to reason about and tackle the non-stationarity of multi-agent traffic environments by both acquiring feature representation distributions and deriving variational lower bounds (i.e., objective functions); And then, based on the designed CI model, we propose a CI-MA algorithm, in which the feature representations are acquired from the non-stationarity of multi-agent traffic environments at both task level and timestep level, the acquired feature representations are used to produce cooperative traffic-signal policies and Q-values for multiple agents; Finally the corresponding objective functions optimize the whole algorithm from both causal inference and multi-agent reinforcement learning. Experiments are conducted in different non-stationary multi-agent traffic environments. Results show that CI-MA algorithm outperforms other state-of-the-art algorithms, and demonstrate that the proposed algorithm trained in synthetic-traffic environments can be effectively transferred to both synthetic- and real-traffic environments with non-stationarity.