Adaptive control of discounted Markov decision chains

Adaptive control of discounted Markov decision chains
复制标题

贴现马尔可夫决策链的自适应控制

DOI:
10.1007/bf00938426
复制
发表时间:
1985
影响因子:
1.9
通讯作者:
Steven I. Marcus
Steven I. Marcus
中科院分区:
数学3区
文献类型:
--
作者:
O. Hernández;Steven I. Marcus

文献摘要

被引文献

相似文献

本文考虑依赖于未知参数的折扣报酬有限状态马尔可夫决策过程。一种受Federgruen和Schweitzer的非定常值迭代格式启发的自适应策略(参考文献)。1)提出了一种新的方法。将这一政策与朔伊布勒·L最近提出的估计与控制原则作了简要的比较。4)。
In this paper, we consider discounted-reward finite-state Markov decision processes which depend on unknown parameters. An adaptive policy inspired by the nonstationary value iteration scheme of Federgruen and Schweitzer (Ref. 1) is proposed. This policy is briefly compared with the principle of estimation and control recently obtained by Schäl (Ref. 4).