Uncertainty Quantification and Exploration for Reinforcement Learning

Uncertainty Quantification and Exploration for Reinforcement Learning
复制标题

DOI:
10.1287/opre.2023.2436
复制
发表时间:
2023-03
影响因子:
2.7
通讯作者:
Yi Zhu;Jing Dong;Henry Lam
Yi Zhu;Jing Dong;Henry Lam
中科院分区:
管理学3区
文献类型:
--
作者:
Yi Zhu;Jing Dong;Henry Lam

文献摘要

相似文献

在统计推断中,大样本行为和置信区间构造是评估估计量相对于数据噪声的误差和可靠性的基础。在论文“Uncertainty Quantification and Exploration for Reinforcement Learning”中,Dong,Lam和Zhu研究了经典强化学习环境中的大样本行为。他们推导出适当的大样本渐近分布的状态-动作值函数(Q值)和最佳值函数估计时,数据收集的基础马尔可夫链。这使人们能够评估不同决策中表现的自信程度。严格的不确定性量化也有利于开发一个纯粹的勘探政策,最大限度地提高最坏情况下的估计Q值之间的相对差异(均方差的方差比)。这种探索策略旨在收集信息丰富的训练数据,以最大限度地提高学习最优奖励收集策略的概率,并取得了良好的经验性能。
Quantify the uncertainty to decide and explore better In statistical inference, large-sample behavior and confidence interval construction are fundamental in assessing the error and reliability of estimated quantities with respect to the data noises. In the paper “Uncertainty Quantification and Exploration for Reinforcement Learning”, Dong, Lam, and Zhu study the large sample behavior in the classic setting of reinforcement learning. They derive appropriate large-sample asymptotic distributions for the state-action value function (Q-value) and optimal value function estimations when data are collected from the underlying Markov chain. This allows one to evaluate the assertiveness of performances among different decisions. The tight uncertainty quantification also facilitates the development of a pure exploration policy by maximizing the worst-case relative discrepancy among the estimated Q-values (ratio of the mean squared difference to the variance). This exploration policy aims to collect informative training data to maximize the probability of learning the optimal reward collecting policy, and it achieves good empirical performance.