Inference and dynamic decision-making for deteriorating systems with probabilistic dependencies through Bayesian networks and deep reinforcement learning

Inference and dynamic decision-making for deteriorating systems with probabilistic dependencies through Bayesian networks and deep reinforcement learning
复制标题

DOI:
10.1016/j.ress.2023.109144
复制
发表时间:
2022-09
期刊:
Reliab. Eng. Syst. Saf.
影响因子:
--
通讯作者:
P. G. Morato;C. Andriotis;K. Papakonstantinou;P. Rigo
P. G. Morato;C. Andriotis;K. Papakonstantinou;P. Rigo
中科院分区:
其他
文献类型:
--
作者:
P. G. Morato;C. Andriotis;K. Papakonstantinou;P. Rigo

文献摘要

被引文献

相似文献

在现代工程、环境和社会问题的背景下,人们对能够确定土木工程系统合理管理策略的方法的需求越来越大,从而最大限度地降低结构失效风险,同时最佳规划检查和维护(I&M)流程。大多数可用的方法简化的I&M决策问题的组件级,通常假设组件之间的统计,结构或成本独立性,由于与联合系统级状态描述下的全局优化方法的计算复杂性。在本文中,我们提出了一个有效的算法框架的推理和决策的不确定性下的工程系统暴露在不断恶化的环境,直接在系统级提供最佳的管理策略。在我们的方法中,决策问题被制定为一个因子化的部分可观察马尔可夫决策过程,其动态编码在贝叶斯网络条件结构。该方法可以处理环境下的平等或一般,不等恶化相关性组件之间,通过高斯层次结构和动态贝叶斯网络,解耦原来的联合系统状态空间的组件网络条件下共享的随机变量。在策略优化方面,我们采用了深度分散的多智能体演员-评论家(DDMAC)强化学习方法,其中策略由评论家网络指导的演员神经网络近似。通过包括在模拟环境中的恶化依赖,并制定成本模型在系统层面上,DDMAC政策本质上考虑潜在的系统效应。这是证明通过数值实验进行了9出10系统和钢框架下的疲劳退化。结果表明,DDMAC政策相比,国家的最先进的启发式方法提供了巨大的好处。DDMAC策略对系统影响的内在考虑也被解释为基于学习的策略。
In the context of modern engineering, environmental, and societal concerns, there is an increasing demand for methods able to identify rational management strategies for civil engineering systems, minimizing structural failure risks while optimally planning inspection and maintenance (I&M) processes. Most available methods simplify the I&M decision problem to the component level, often assuming statistical, structural, or cost independence among components, due to the computational complexity associated with global optimization methodologies under joint system-level state descriptions. In this paper, we propose an efficient algorithmic framework for inference and decision-making under uncertainty for engineering systems exposed to deteriorating environments, providing optimal management strategies directly at the system level. In our approach, the decision problem is formulated as a factored partially observable Markov decision process, whose dynamics are encoded in Bayesian network conditional structures. The methodology can handle environments under equal or general, unequal deterioration correlations among components, through Gaussian hierarchical structures and dynamic Bayesian networks, decoupling the originally joint system state space to component networks conditional on shared random variables. In terms of policy optimization, we adopt a deep decentralized multi-agent actor-critic (DDMAC) reinforcement learning approach, in which the policies are approximated by actor neural networks guided by a critic network. By including deterioration dependence in the simulated environment, and by formulating the cost model at the system level, DDMAC policies intrinsically consider the underlying system-effects. This is demonstrated through numerical experiments conducted for both a 9-out-of-10 system and a steel frame under fatigue deterioration. Results demonstrate that DDMAC policies offer substantial benefits when compared to state-of-the-art heuristic approaches. The inherent consideration of system-effects by DDMAC strategies is also interpreted based on the learned policies.