Approximate Dynamic Programming by Practical Examples
Approximate Dynamic Programming by Practical Examples
复制标题
实例近似动态规划
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
A. P. Rivera
中科院分区:
文献类型:
--
作者:
M. Mes;A. P. Rivera
Computing the exact solution of an MDP model is generally difficult and possibly intractable for realistically sized problem instances. A powerful technique to solve the large scale discrete time multistage stochastic control processes is Approximate Dynamic Programming (ADP). Although ADP is used as an umbrella term for a broad spectrum of methods to approximate the optimal solution of MDPs, the common denominator is typically to combine optimization with simulation, use approximations of the optimal values of the Bellman’s equations, and use approximate policies. This chapter aims to present and illustrate the basics of these steps by a number of practical and instructive examples. We use three examples (1) to explain the basics of ADP, relying on value iteration with an approximation of the value functions, (2) to provide insight into implementation issues, and (3) to provide test cases for the reader to validate its own ADP implementations.