An Approximately Optimal Relative Value Learning Algorithm for Averaged MDPs with Continuous States and Actions

An Approximately Optimal Relative Value Learning Algorithm for Averaged MDPs with Continuous States and Actions
复制标题

具有连续状态和动作的平均 MDP 的近似最优相对值学习算法

DOI:
10.1109/allerton.2019.8919719
复制
发表时间:
2019
期刊:
and Computing (Allerton
影响因子:
--
通讯作者:
Jain, Rahul
Jain, Rahul
中科院分区:
--
文献类型:
--
作者:
Sharma, Hiteshi;Jain, Rahul

文献摘要

参考文献

相似文献

DOI: 10.1109/cdc40024.2019.9029308
发表时间: 2019-12
期刊: 2019 IEEE 58th Conference on Decision and Control (CDC)
影响因子: --
作者:
Hiteshi Sharma;R. Jain;W. Haskell
通讯作者: Hiteshi Sharma;R. Jain;W. Haskell
DOI: 10.1109/tit.1978.1055865
发表时间: 1978
期刊: IEEE Trans. Inf. Theory
影响因子: --
作者:
L. Devroye
通讯作者: L. Devroye
DOI: 10.1007/bf01585710
发表时间: 1992-02
影响因子: 2.7
作者:
Z. Zabinsky;Robert L. Smith
通讯作者: Z. Zabinsky;Robert L. Smith
平均奖励马尔可夫决策过程中状态聚合的伪计量学
DOI: 10.1007/978-3-540-75225-7_30
发表时间: 2007
期刊: 2019 18th European Control Conference (ECC)
影响因子: --
作者:
R. Ortner
通讯作者: R. Ortner
DOI: 10.23919/ecc.2019.8795982
发表时间: 2019
期刊: 2019 18th European Control Conference (ECC
影响因子: --
作者:
Sharma, Hiteshi;Jain, Rahul;Gupta, Abhishek
通讯作者: Gupta, Abhishek