Model-free reinforcement learning with model-based safe exploration: Optimizing adaptive recovery process of infrastructure systems
Model-free reinforcement learning with model-based safe exploration: Optimizing adaptive recovery process of infrastructure systems
复制标题
基于模型的安全探索的无模型强化学习:优化基础设施系统的自适应恢复过程
DOI:
10.1016/j.strusafe.2019.04.003
复制
发表时间:
2019
影响因子:
5.8
通讯作者:
Pozzi, Matteo
中科院分区:
文献类型:
--
作者:
Memarzadeh, Milad;Pozzi, Matteo
Extreme events represent not only some of the most damaging events in our society and environment, but also the most difficult to predict. Model-based predictions of the disruptions induced by extreme events on urban infrastructure systems are often unreliable, as these events are unlikely by their very definition. Specifically, characterizing the effect of such disruptions to the urban infrastructure using a parameterized model is a difficult task. On the other hand, model-free approaches based on recent advancements in reinforcement learning can model the complex dynamics of urban society and infrastructure under the risk of extreme events explicitly without relying on any specific physics-based mechanism. However, these approaches usually require performing random exploration of the effects of management actions on the system (typically in the post-event situation) to allow for an acceptable approximation to the optimal management policy. When dealing with costly infrastructure systems and important communities, this random exploration can be unacceptable and risky. In this paper, we propose a method called Safe Q-learning, which is a model-free reinforcement learning approach with addition of a model-based safe exploration for near-optimal management of infrastructure system pre-event and their recovery post-event. Our method requires the decision-maker to model the structure of the state space of the problem, and a suitable equilibrium of the system (optimum functionality pre-event). This information is usually available for urban systems, as they spend long time in optimum equilibrium before the occurrence of such events. We show on several examples of infrastructure management how the proposed approach is able to achieve near-optimal performance without the risk due to random exploration.
登录
查看更多内容
影响因子:
3.3
作者:
Bristow, David N.;Hay, Alexander H.
通讯作者:
Hay, Alexander H.
DOI:
--
发表时间:
2017
期刊:
6–10 August 2017
影响因子:
--
作者:
Pozzi, M.;Memarzadeh, M.
通讯作者:
Memarzadeh, M.
影响因子:
5.8
作者:
J. Luque;D. Štraub
通讯作者:
J. Luque;D. Štraub
DOI:
10.1016/0968-090x(93)90021-7
发表时间:
1993-03
影响因子:
8.3
作者:
S. Madanat
通讯作者:
S. Madanat
DOI:
10.1016/j.ress.2016.04.016
发表时间:
2016
期刊:
Reliab. Eng. Syst. Saf.
影响因子:
--
作者:
Milad Memarzadeh;M. Pozzi;J. Z. Kolter
通讯作者:
J. Z. Kolter