Adaptive control of discounted Markov decision chains
Adaptive control of discounted Markov decision chains
复制标题
贴现马尔可夫决策链的自适应控制
DOI:
10.1007/bf00938426
复制
发表时间:
1985
影响因子:
1.9
通讯作者:
Steven I. Marcus
中科院分区:
文献类型:
--
作者:
O. Hernández;Steven I. Marcus
In this paper, we consider discounted-reward finite-state Markov decision processes which depend on unknown parameters. An adaptive policy inspired by the nonstationary value iteration scheme of Federgruen and Schweitzer (Ref. 1) is proposed. This policy is briefly compared with the principle of estimation and control recently obtained by Schäl (Ref. 4).