Uncertainty-aware Energy Management of Extended Range Electric Delivery Vehicles with Bayesian Ensemble

Uncertainty-aware Energy Management of Extended Range Electric Delivery Vehicles with Bayesian Ensemble
复制标题

DOI:
10.1109/iv47402.2020.9304826
复制
发表时间:
2020-10
期刊:
2020 IEEE Intelligent Vehicles Symposium (IV)
影响因子:
--
通讯作者:
Pengyue Wang;Yan Li;S. Shekhar;W. Northrop
Pengyue Wang;Yan Li;S. Shekhar;W. Northrop
中科院分区:
其他
文献类型:
--
作者:
Pengyue Wang;Yan Li;S. Shekhar;W. Northrop

文献摘要

相似文献

近年来,深度强化学习(DRL)算法在智能交通系统(ITS)领域得到了广泛的研究和应用。DRL智能体大多使用仿真生成的转换对和交互轨迹进行训练,并且它们可以在熟悉的输入状态下实现令人满意或接近最优的性能。然而,对于状态空间中相对较少访问甚至未访问的区域,不能保证智能体能够良好地执行。不幸的是,在现实世界的问题中,新的条件是不可避免的,并且在真实的数据和模拟数据之间总是存在差距。因此,要在现实世界的交通系统中实现DRL算法,我们不仅要训练智能体学习将状态映射到动作的策略,还要学习与每个动作相关的模型不确定性。在这项研究中,我们采用贝叶斯集成的方法来训练一组代理人施加的多样性的能源管理系统的交付车辆。在合奏的代理人同意熟悉的国家,但不熟悉或新的国家表现出不同的结果。这种不确定性估计有助于实现可解释的后处理模块,从而确保在高不确定性条件下的稳健和安全操作。
In recent years, deep reinforcement learning (DRL) algorithms have been widely studied and utilized in the area of Intelligent Transportation Systems (ITS). DRL agents are mostly trained with transition pairs and interaction trajectories generated from simulation, and they can achieve satisfying or near optimal performances under familiar input states. However, for relative rare visited or even unvisited regions in the state space, there is no guarantee that the agent could perform well. Unfortunately, novel conditions are inevitable in real-world problems and there is always a gap between the real data and simulated data. Therefore, to implement DRL algorithms in real-world transportation systems, we should not only train the agent learn a policy that maps states to actions, but also the model uncertainty associated with each action. In this study, we adapt the method of Bayesian ensemble to train a group of agents with imposed diversity for an energy management system of a delivery vehicle. The agents in the ensemble agree well on familiar states but show diverse results on unfamiliar or novel states. This uncertainty estimation facilitates the implementation of interpretable postprocessing modules which can ensure robust and safe operations under high uncertainty conditions.