Delayed Feedback in Episodic Reinforcement Learning

Delayed Feedback in Episodic Reinforcement Learning
复制标题

情景强化学习中的延迟反馈

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
S. Filippi
S. Filippi
中科院分区:
--
文献类型:
--
作者:
Benjamin Howson;Ciara Pike;S. Filippi

文献摘要

被引文献

相似文献

有许多经过证明有效的情景强化学习算法。然而,这些算法是在这样的假设下构建的:与每个情节相关的状态、动作和奖励序列立即到达,允许在每次与环境交互后更新策略。这种假设在实践中通常是不切实际的,特别是在医疗保健和在线推荐等领域。在本文中,我们研究了延迟反馈对情景强化学习中后悔最小化的几种已证明有效的算法的影响。首先,我们考虑在有新反馈后立即更新政策。使用这种更新方案,我们表明遗憾会随着涉及状态数量、动作、情节长度和预期延迟的附加项而增加。该附加项根据选择的乐观算法而变化。我们还表明,不那么频繁地更新政策可能会导致遗憾对延迟的依赖程度增加。
There are many provably efficient algorithms for episodic reinforcement learning. However, these algorithms are built under the assumption that the sequences of states, actions and rewards associated with each episode arrive immediately, allowing policy updates after every interaction with the environment. This assumption is often unrealistic in practice, particularly in areas such as healthcare and online recommendation. In this paper, we study the impact of delayed feedback on several provably efficient algorithms for regret minimisation in episodic reinforcement learning. Firstly, we consider updating the policy as soon as new feedback becomes available. Using this updating scheme, we show that the regret increases by an additive term involving the number of states, actions, episode length and the expected delay. This additive term changes depending on the optimistic algorithm of choice. We also show that updating the policy less frequently can lead to an improved dependency of the regret on the delays.