Stochastic optimal well control in subsurface reservoirs using reinforcement learning

Stochastic optimal well control in subsurface reservoirs using reinforcement learning
复制标题

DOI:
10.1016/j.engappai.2022.105106
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Dixit;A. Elsheikh
A. Dixit;A. Elsheikh
中科院分区:
其他
文献类型:
--
作者:
A. Dixit;A. Elsheikh

文献摘要

相似文献

我们提出了一个无模型强化学习(RL)框架的案例研究,用于解决预定义参数不确定性分布和部分可观测系统的随机最优控制。我们专注于鲁棒最优井控问题,这是地下油藏管理领域深入研究活动的一个主题。对于这个问题,系统被部分观察,因为数据仅在井位可用。此外,由于可用现场数据的稀疏性,模型参数具有高度不确定性。原则上,强化学习算法能够学习最佳行动策略(从状态到行动的映射),以最大化数值奖励信号。在深度强化学习中,从状态到动作的映射使用深度神经网络进行参数化。在鲁棒最优井控问题的强化学习公式中,状态由井位处的饱和度和压力值表示,而动作表示控制通过井的流量的阀门开度。数值奖励是指总波及效率,不确定模型参数是地下渗透率场。通过引入域随机化方案来处理模型参数的不确定性,该方案利用对其不确定性分布的聚类分析。我们使用两种最先进的 RL 算法(近端策略优化 (PPO) 和优势行动者批评家 (A2C))在代表渗透率场的两种不同不确定性分布的两个地下流动测试案例上给出数值结果。结果与使用差分进化算法获得的优化结果进行了基准比较。此外,我们通过评估从训练过程中未使用的参数不确定性分布中提取的未见样本上学习到的控制策略,证明了所提出的 RL 使用的鲁棒性。
We present a case study of model-free reinforcement learning (RL) framework to solve stochastic optimal control for a predefined parameter uncertainty distribution and partially observable system. We focus on robust optimal well control problem which is a subject of intensive research activities in the field of subsurface reservoir management. For this problem, the system is partially observed since the data is only available at well locations. Furthermore, the model parameters are highly uncertain due to sparsity of available field data. In principle, RL algorithms are capable of learning optimal action policies – a map from states to actions – to maximize a numerical reward signal. In deep RL, this mapping from state to action is parameterized using a deep neural network. In the RL formulation of the robust optimal well control problem, the states are represented by saturation and pressure values at well locations while the actions represent the valve openings controlling the flow through wells. The numerical reward refers to the total sweep efficiency and the uncertain model parameter is the subsurface permeability field. The model parameter uncertainties are handled by introducing a domain randomization scheme that exploits cluster analysis on its uncertainty distribution. We present numerical results using two state-of-the-art RL algorithms, proximal policy optimization (PPO) and advantage actor–critic (A2C), on two subsurface flow test cases representing two distinct uncertainty distributions of permeability field. The results were benchmarked against optimization results obtained using differential evolution algorithm. Furthermore, we demonstrate the robustness of the proposed use of RL by evaluating the learned control policy on unseen samples drawn from the parameter uncertainty distribution that were not used during the training process.