A novel nested stochastic dynamic programming (nSDP) and nested reinforcement learning (nRL) algorithm for multipurpose reservoir optimization

A novel nested stochastic dynamic programming (nSDP) and nested reinforcement learning (nRL) algorithm for multipurpose reservoir optimization
复制标题

一种用于多用途油藏优化的新型嵌套随机动态规划(nSDP)和嵌套强化学习(nRL)算法

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
D. Solomatine
D. Solomatine
中科院分区:
--
文献类型:
--
作者:
Blagoj Delipetrev;A. Jonoski;D. Solomatine

文献摘要

被引文献

相似文献

本文提出了两种新的多目标水库优化算法:嵌套随机动态规划算法(NSDP)和嵌套强化学习算法(NRL)。这两种算法都被构建为两种算法的组合;在nSDP的情况下,它是(1)随机动态规划(SDP)和(2)嵌套最优分配算法(NOAA),而在NRL的情况下,它是(1)强化学习(RL)和(2)NOAA。NOAA采用线性和非线性两种优化方法。主要的新颖思想是在每个SDP和RL状态转换处包括NOAA,这降低了初始问题的维度,并缓解了维度诅咒。NSDP和NRL都能在不增加计算开销和算法复杂度的情况下解决多目标优化问题,并能处理密集和不规则变量的离散化。这两个算法是在位于马其顿共和国的Knezevo水库上用Java编写的原型应用程序。将nSDP和NRL最优水库策略与嵌套动态规划策略进行了比较,总体结论是NRL比nSDP更有效,但其复杂性明显高于nSDP。
In this article we present two novel multipurpose reservoir optimization algorithms named nested stochastic dynamic programming (nSDP) and nested reinforcement learning (nRL). Both algorithms are built as a combination of two algorithms; in the nSDP case it is (1) stochastic dynamic programming (SDP) and (2) nested optimal allocation algorithm (nOAA) and in the nRL case it is (1) reinforcement learning (RL) and (2) nOAA. The nOAA is implemented with linear and non-linear optimization. The main novel idea is to include a nOAA at each SDP and RL state transition, that decreases starting problem dimension and alleviates curse of dimensionality. Both nSDP and nRL can solve multi-objective optimization problems without significant computational expenses and algorithm complexity and can handle dense and irregular variable discretization. The two algorithms were coded in Java as a prototype application and on the Knezevo reservoir, located in the Republic of Macedonia. The nSDP and nRL optimal reservoir policies were compared with nested dynamic programming policies, and overall conclusion is that nRL is more powerful, but significantly more complex than nSDP.