Markov decision Processes with fractional costs

Markov decision Processes with fractional costs
复制标题

具有分数成本的马尔可夫决策过程

DOI:
--
复制
发表时间:
2005
影响因子:
6.8
通讯作者:
B. Krogh
B. Krogh
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhiyuan Ren;B. Krogh

文献摘要

被引文献

相似文献

某些用于构建嵌入式马尔可夫决策过程(MDP)的方法导致性能度量是两个长期平均值的比率。对于这种具有有限状态和动作空间的MDP,在遍历性假设下,本文提出了基于策略迭代、线性规划、值迭代和Q学习的最优策略计算算法。
Certain methods for constructing embedded Markov decision processes (MDPs) lead to performance measures that are the ratio of two long-run averages. For such MDPs with finite state and action spaces and under an ergodicity assumption, this note presents algorithms for computing optimal policies based on policy iterations, linear programming, value iterations and Q-learning.