Off-Policy Confidence Interval Estimation with Confounded Markov Decision Process
Off-Policy Confidence Interval Estimation with Confounded Markov Decision Process
复制标题
混杂马尔可夫决策过程的离策略置信区间估计
DOI:
10.1080/01621459.2022.2110878
复制
发表时间:
2022
影响因子:
3.7
通讯作者:
Shi C
中科院分区:
文献类型:
--
作者:
Shi C
This article is concerned with constructing a confidence interval for a target policy’s value offline based on a pre-collected observational data in infinite horizon settings. Most of the existing works assume no unmeasured variables exist that confound the observed actions. This assumption, however, is likely to be violated in real applications such as healthcare and technological industries. In this article, we show that with some auxiliary variables that mediate the effect of actions on the system dynamics, the target policy’s value is identifiable in a confounded Markov decision process. Based on this result, we develop an efficient off-policy value estimator that is robust to potential model misspecification and provide rigorous uncertainty quantification. Our method is justified by theoretical results, simulated and real datasets obtained from ridesharing companies. A Python implementation of the proposed procedure is available athttps://github.com/Mamba413/cope.
登录
查看更多内容
DOI:
--
发表时间:
2007
期刊:
Conference on Uncertainty in Artificial Intelligence
影响因子:
--
作者:
Michael P. Holmes;Alexander G. Gray;C. Isbell
通讯作者:
C. Isbell
DOI:
--
发表时间:
2019
期刊:
FAT
影响因子:
--
作者:
Alexander Luedtke;M. J. van der Laan
通讯作者:
M. J. van der Laan
DOI:
--
发表时间:
2020
期刊:
IEEE/RJS International Conference on Intelligent RObots and Systems
影响因子:
--
作者:
Chengxi Li;Stanley H. Chan;Yi
通讯作者:
Yi
影响因子:
2.7
作者:
Zhang B;Tsiatis AA;Laber EB;Davidian M
通讯作者:
Davidian M
影响因子:
2
作者:
Chernofsky,Ariel;Bosch,RonaldJ;Lok,JudithJ
通讯作者:
Lok,JudithJ