Performance Loss Bounds for Approximate Value Iteration with State Aggregation

Performance Loss Bounds for Approximate Value Iteration with State Aggregation
复制标题

DOI:
10.1287/moor.1060.0188
复制
发表时间:
2006-02
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
Benjamin Van Roy
Benjamin Van Roy
中科院分区:
其他
文献类型:
--
作者:
Benjamin Van Roy

文献摘要

被引文献

相似文献

我们考虑近似值迭代与参数化的逼近器,其中状态空间被划分,并且每个分区上的最优成本函数由常数近似。我们建立性能损失界限的政策来自近似与不动点。这些界限确定了使用适当策略的不变分布作为投影权重的好处。这种投影加权与时间差异学习所做的事情有关。我们的分析还导致了第一个性能损失的近似值迭代的平均成本目标。
We consider approximate value iteration with a parameterized approximator in which the state space is partitioned and the optimal cost-to-go function over each partition is approximated by a constant. We establish performance loss bounds for policies derived from approximations associated with fixed points. These bounds identify benefits to using invariant distributions of appropriate policies as projection weights. Such projection weighting relates to what is done by temporal-difference learning. Our analysis also leads to the first performance loss bound for approximate value iteration with an average-cost objective.