Model-free reinforcement learning with model-based safe exploration: Optimizing adaptive recovery process of infrastructure systems

Model-free reinforcement learning with model-based safe exploration: Optimizing adaptive recovery process of infrastructure systems
复制标题

基于模型的安全探索的无模型强化学习:优化基础设施系统的自适应恢复过程

DOI:
10.1016/j.strusafe.2019.04.003
复制
发表时间:
2019
期刊:
影响因子:
5.8
通讯作者:
Pozzi, Matteo
Pozzi, Matteo
中科院分区:
工程技术1区
文献类型:
--
作者:
Memarzadeh, Milad;Pozzi, Matteo

文献摘要

参考文献

被引文献

相似文献

极端事件不仅代表了我们社会和环境中一些最具破坏性的事件,也是最难以预测的事件。对极端事件对城市基础设施系统造成的破坏,基于模型的预测往往是不可靠的,因为这些事件本身就不太可能发生。具体地说,使用参数化模型来表征这种干扰对城市基础设施的影响是一项艰巨的任务。另一方面,基于强化学习的最新进展的无模型方法可以明确地模拟极端事件风险下城市社会和基础设施的复杂动态,而不依赖于任何特定的基于物理的机制。然而,这些方法通常需要对管理操作对系统的影响执行随机探索(通常在事后情况下),以允许对最佳管理策略进行可接受的近似。在处理昂贵的基础设施系统和重要社区时,这种随机探索可能是不可接受的,也是有风险的。在本文中,我们提出了一种称为安全Q-学习的方法,它是一种无模型的强化学习方法,并增加了一种基于模型的安全探索,用于基础设施系统的事前和事后恢复的近最佳管理。我们的方法需要决策者对问题的状态空间的结构进行建模,以及系统的适当平衡(事件前的最优功能)。这种信息通常适用于城市系统,因为在此类事件发生之前,城市系统需要花费很长时间处于最佳平衡状态。我们在几个基础设施管理的例子中展示了所提出的方法如何能够获得接近最优的性能,而不会因随机探索而产生风险。
Extreme events represent not only some of the most damaging events in our society and environment, but also the most difficult to predict. Model-based predictions of the disruptions induced by extreme events on urban infrastructure systems are often unreliable, as these events are unlikely by their very definition. Specifically, characterizing the effect of such disruptions to the urban infrastructure using a parameterized model is a difficult task. On the other hand, model-free approaches based on recent advancements in reinforcement learning can model the complex dynamics of urban society and infrastructure under the risk of extreme events explicitly without relying on any specific physics-based mechanism. However, these approaches usually require performing random exploration of the effects of management actions on the system (typically in the post-event situation) to allow for an acceptable approximation to the optimal management policy. When dealing with costly infrastructure systems and important communities, this random exploration can be unacceptable and risky. In this paper, we propose a method called Safe Q-learning, which is a model-free reinforcement learning approach with addition of a model-based safe exploration for near-optimal management of infrastructure system pre-event and their recovery post-event. Our method requires the decision-maker to model the structure of the state space of the problem, and a suitable equilibrium of the system (optimum functionality pre-event). This information is usually available for urban systems, as they spend long time in optimum equilibrium before the occurrence of such events. We show on several examples of infrastructure management how the proposed approach is able to achieve near-optimal performance without the risk due to random exploration.
DOI: 10.1061/(asce)is.1943-555x.0000338
发表时间: 2017-09-01
影响因子: 3.3
作者:
Bristow, David N.;Hay, Alexander H.
通讯作者: Hay, Alexander H.
DOI: --
发表时间: 2017
期刊: 6–10 August 2017
影响因子: --
作者:
Pozzi, M.;Memarzadeh, M.
通讯作者: Memarzadeh, M.
DOI: 10.1016/j.strusafe.2018.08.002
发表时间: 2019
期刊: Structural Safety
影响因子: 5.8
作者:
J. Luque;D. Štraub
通讯作者: J. Luque;D. Štraub
DOI: 10.1016/0968-090x(93)90021-7
发表时间: 1993-03
影响因子: 8.3
作者:
S. Madanat
通讯作者: S. Madanat
具有相似组件的系统的分层建模:自适应监视和控制的框架
DOI: 10.1016/j.ress.2016.04.016
发表时间: 2016
期刊: Reliab. Eng. Syst. Saf.
影响因子: --
作者:
Milad Memarzadeh;M. Pozzi;J. Z. Kolter
通讯作者: J. Z. Kolter