Simulation of pedestrian evacuation with reinforcement learning based on a dynamic scanning algorithm

Simulation of pedestrian evacuation with reinforcement learning based on a dynamic scanning algorithm
复制标题

DOI:
10.1016/j.physa.2023.129011
复制
发表时间:
2023-06
期刊:
Physica A: Statistical Mechanics and its Applications
影响因子:
--
通讯作者:
Zhongyi Huang;Rong Liang;Yao Xiao;Zhiming Fang;Xiaolian Li;Rui Ye
Zhongyi Huang;Rong Liang;Yao Xiao;Zhiming Fang;Xiaolian Li;Rui Ye
中科院分区:
其他
文献类型:
--
作者:
Zhongyi Huang;Rong Liang;Yao Xiao;Zhiming Fang;Xiaolian Li;Rui Ye

文献摘要

相似文献

人类主要根据视觉信息来计划他们的动作。然而,现有疏散模型中的代理很少通过视觉信息来感知环境。为了获取离散场中个体的视觉特征,提出了动态扫描算法(DSA)。 DSA 引入了激光雷达的射线扫描,代理释放“激光”来检测与之相交的最近的物体。通过预先存储射线穿过的网格,显着提高了DSA的效率。使用扫描结果作为输入,基于双 Q 学习(DDQN)的深度强化学习开发了疏散模型。首先对DSA的参数进行标定,推荐一组效率与精度平衡良好的参数。此外,还复制了基本图来校准 DDQN 中的奖励值。最后,利用标定的参数研究模型中的轨迹和行为。结果表明,运动轨迹受到可见距离和奖励值的影响,部分效果与实验结果一致。此外,在仿真中观察出口选择行为和车道形成行为,无需引入任何特殊设计规则。 DSA提供了一种获取离散场第一人称环境信息的新方法,基于DSA&DDQN的模型定义了一种新的基于射线扫描的疏散建模方案。
Humans plan their movements mainly based on visual information. However, agents in few existing evacuation models perceive the environment by using visual information. To obtain the visual features of an individual in a discrete field, a Dynamic Scanning Algorithm (DSA) is proposed. DSA introduces the ray-scanning of LIDAR, a ”laser” is released by an agent to detect the nearest object that intersects it. By pre-storing the grids crossed by the rays, the efficiency of DSA is significantly improved. Using the scan results as inputs, an evacuation model has been developed based on the Deep Reinforcement Learning with Double Q-learning (DDQN). The parameters of DSA are calibrated at first, and a group of parameters with a good balance between efficiency and accuracy are recommended. Furthermore, the fundamental diagram is reproduced to calibrate the reward values in DDQN. At last, trajectories and behaviors in the model are studied by using the calibrated parameters. Results show that the movement trajectories are affected by visible distances and reward values, and some effectiveness are consistent with that in experiments. Besides, the exit selection behavior and the lane formation behavior are observed in simulation without introducing any special designed rules. DSA provides a new method to obtain the first-person environmental information in discrete field, and the DSA&DDQN based model defines a new ray-scan-based evacuation modeling scheme.