Robust Optimal Well Control using an Adaptive Multigrid Reinforcement Learning Framework

Robust Optimal Well Control using an Adaptive Multigrid Reinforcement Learning Framework
复制标题

DOI:
10.1007/s11004-022-10033-x
复制
发表时间:
2022-11-04
影响因子:
2.6
通讯作者:
Elsheikh, Ahmed H.
Elsheikh, Ahmed H.
中科院分区:
地球科学3区
文献类型:
--
作者:
Dixit, Atish;Elsheikh, Ahmed H.

文献摘要

被引文献

相似文献

强化学习(RL)是一种很有前途的工具,用于解决鲁棒最优井控问题的模型参数是高度不确定的,在实践中的系统是部分可观测的。然而,鲁棒控制策略的RL通常依赖于执行大量的仿真。对于计算密集型模拟的情况,这很容易变得计算上难以处理。为了解决这一瓶颈,自适应多重网格RL框架介绍了几何多重网格方法在迭代数值算法中使用的原则的启发。RL控制策略最初是使用计算效率高的低保真度模拟与基本偏微分方程(PDE)的粗网格离散化来学习的。随后,模拟保真度以自适应的方式增加到最高保真度的模拟,对应于模型域的最精细的离散化。所提出的框架使用最先进的,无模型的基于策略的强化学习算法,即邻近策略优化算法。结果显示两个案例研究的强大的最优井控制问题,这是从SPE-10模型2基准案例研究的启发。使用所提出的框架可以观察到计算效率的显著提高,节省了其单个细网格对应物的计算成本的60-70%。
Reinforcement learning (RL) is a promising tool for solving robust optimal well control problems where the model parameters are highly uncertain and the system is partially observable in practice. However, the RL of robust control policies often relies on performing a large number of simulations. This could easily become computationally intractable for cases with computationally intensive simulations. To address this bottleneck, an adaptive multigrid RL framework is introduced which is inspired by principles of geometric multigrid methods used in iterative numerical algorithms. RL control policies are initially learned using computationally efficient low-fidelity simulations with coarse grid discretization of the underlying partial differential equations (PDEs). Subsequently, the simulation fidelity is increased in an adaptive manner towards the highest fidelity simulation that corresponds to the finest discretization of the model domain. The proposed framework is demonstrated using a state-of-the-art, model-free policy-based RL algorithm, namely the proximal policy optimization algorithm. Results are shown for two case studies of robust optimal well control problems, which are inspired from SPE-10 model 2 benchmark case studies. Prominent gains in computational efficiency are observed using the proposed framework, saving around 60-70% of the computational cost of its single fine-grid counterpart.