Managing engineering systems with large state and action spaces through deep reinforcement learning

Managing engineering systems with large state and action spaces through deep reinforcement learning
复制标题

DOI:
10.1016/j.ress.2019.04.036
复制
发表时间:
2019-11-01
影响因子:
8.1
通讯作者:
Papakonstantinou, K. G.
Papakonstantinou, K. G.
中科院分区:
工程技术1区
文献类型:
--
作者:
Andriotis, C. P.;Papakonstantinou, K. G.

文献摘要

被引文献

相似文献

可以使用马尔可夫决策过程 (MDP) 或部分可观察 MDP (POMDP) 有效地制定工程系统管理决策。典型的 MDP/POMDP 解决方案过程利用有关环境的离线知识,并为具有易于处理的状态和操作空间的相对较小的系统提供详细的策略。然而,在大型多组件系统中,这些空间的维度很容易爆炸,因为系统状态和动作随着组件数量呈指数级增长,而整个系统的环境动态很难明确描述,并且通常只能通过计算成本昂贵的数值模拟器来访问。在这项工作中,为了解决这些问题,引入了集成的深度强化学习(DRL)框架。开发了深度集中式多智能体行为评论家 (DCMAC),这是一种离策略行为评论家 DRL 算法,可直接探测底层 MDP/POMDP 的状态/置信空间,为在高维空间中运行的大型多组件系统提供高效的生命周期策略。除了具有巨大状态空间的参数化复杂函数的深度网络逼近器之外,DCMAC 还采用系统操作的因式分解表示,从而能够指定个性化的组件和子系统级决策,同时维护整个系统的集中价值函数。 DCMAC 与 Deep Q-Network 和精确解决方案(如果适用)相比效果很好,并且优于包含基于时间、基于条件和定期检查和维护注意事项的优化基线策略。
Decision-making for engineering systems management can be efficiently formulated using Markov Decision Processes (MDPs) or Partially Observable MDPs (POMDPs). Typical MDP/POMDP solution procedures utilize offline knowledge about the environment and provide detailed policies for relatively small systems with tractable state and action spaces. However, in large multi-component systems the dimensions of these spaces easily explode, as system states and actions scale exponentially with the number of components, whereas environment dynamics are difficult to be described explicitly for the entire system and may, often, only be accessible through computationally expensive numerical simulators. In this work, to address these issues, an integrated Deep Reinforcement Learning (DRL) framework is introduced. The Deep Centralized Multi-agent Actor Critic (DCMAC) is developed, an off-policy actor-critic DRL algorithm that directly probes the state/belief space of the underlying MDP/POMDP, providing efficient life-cycle policies for large multi-component systems operating in high-dimensional spaces. Apart from deep network approximators parametrizing complex functions with vast state spaces, DCMAC also adopts a factorized representation of the system actions, thus being able to designate individualized component- and subsystem-level decisions, while maintaining a centralized value function for the entire system. DCMAC compares well against Deep Q-Network and exact solutions, where applicable, and outperforms optimized baseline policies that incorporate time-based, condition-based, and periodic inspection and maintenance considerations.