Approximate Dynamic Programming by Practical Examples

Approximate Dynamic Programming by Practical Examples
复制标题

实例近似动态规划

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
A. P. Rivera
A. P. Rivera
中科院分区:
--
文献类型:
--
作者:
M. Mes;A. P. Rivera

文献摘要

被引文献

相似文献

计算MDP模型的精确解通常是困难的,并且对于实际大小的问题实例可能是棘手的。近似动态规划(ADP)是求解大规模离散时间多阶段随机控制过程的一种强有力的方法。虽然ADP被用作一个总括性术语,用于广泛的方法来近似MDP的最优解,但其共同点通常是将联合收割机优化与仿真相结合,使用Bellman方程的最优值的近似值,并使用近似策略。本章的目的是通过一些实际和有指导意义的例子来介绍和说明这些步骤的基础知识。我们使用三个示例来解释ADP的基础知识,依赖于值迭代和值函数的近似值,(2)提供对实现问题的洞察,以及(3)为读者提供测试用例来验证自己的ADP实现。
Computing the exact solution of an MDP model is generally difficult and possibly intractable for realistically sized problem instances. A powerful technique to solve the large scale discrete time multistage stochastic control processes is Approximate Dynamic Programming (ADP). Although ADP is used as an umbrella term for a broad spectrum of methods to approximate the optimal solution of MDPs, the common denominator is typically to combine optimization with simulation, use approximations of the optimal values of the Bellman’s equations, and use approximate policies. This chapter aims to present and illustrate the basics of these steps by a number of practical and instructive examples. We use three examples (1) to explain the basics of ADP, relying on value iteration with an approximation of the value functions, (2) to provide insight into implementation issues, and (3) to provide test cases for the reader to validate its own ADP implementations.