Q-Learning: Theory and Applications

Q-Learning: Theory and Applications
复制标题

DOI:
10.1146/annurev-statistics-031219-041220
复制
发表时间:
2020-01-01
期刊:
ANNUAL REVIEW OF STATISTICS AND ITS APPLICATION, VOL 7, 2020
影响因子:
--
通讯作者:
Laber, Eric
Laber, Eric
中科院分区:
其他
文献类型:
--
作者:
Clifton, Jesse;Laber, Eric

文献摘要

被引文献

相似文献

Q-学习最初是一种估计无限区间决策问题中最优决策策略的增量式算法,现在指的是一类广泛应用于统计学和人工智能的强化学习方法。在个性化医学的背景下,有限范围Q学习是估计最优治疗策略的主力,即众所周知的治疗方案。无限视野Q学习在不断增长的移动健康领域也变得越来越重要。在计算机科学中,Q-学习方法在游戏和机器人等领域取得了显著的性能。在本文中,我们(A)回顾了计算机科学和统计学中Q-学习的历史,(B)在潜在结果框架内形式化有限时间Q-学习,并讨论了它臭名昭著的推理困难,以及(C)回顾了无限时间Q-学习的变体和在长时间期限决策问题中出现的探索-开发问题。我们最后讨论了在实践中使用Q学习所产生的问题,包括将Q学习与直接搜索方法相结合的论点;序贯、多任务随机试验的样本量考虑;以及将Q学习与基于模型的方法相结合的可能性。
Q-learning, originally an incremental algorithm for estimating an optimal decision strategy in an infinite-horizon decision problem, now refers to a general class of reinforcement learning methods widely used in statistics and artificial intelligence. In the context of personalized medicine, finite-horizon Q-learning is the workhorse for estimating optimal treatment strategies, known as treatment regimes. Infinite-horizon Q-learning is also increasingly relevant in the growing field of mobile health. In computer science, Q-learning methods have achieved remarkable performance in domains such as game-playing and robotics. In this article, we (a) review the history of Q-learning in computer science and statistics, (b) formalize finite-horizon Q-learning within the potential outcomes framework and discuss the inferential difficulties for which it is infamous, and (c) review variants of infinite-horizon Q-learning and the exploration-exploitation problem, which arises in decision problems with a long time horizon. We close by discussing issues arising with the use of Q-learning in practice, including arguments for combining Q-learning with direct-search methods; sample size considerations for sequential, multiple assignment randomized trials; and possibilities for combining Q-learning with model-based methods.