Safely Bridging Offline and Online Reinforcement Learning
Safely Bridging Offline and Online Reinforcement Learning
复制标题
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Wanqiao Xu;Kan Xu;Hamsa Bastani;O. Bastani
中科院分区:
文献类型:
--
作者:
Wanqiao Xu;Kan Xu;Hamsa Bastani;O. Bastani
A key challenge to deploying reinforcement learning in practice is exploring safely. We propose a natural safety property— uniformly outperforming a conservative policy (adaptively estimated from all data observed thus far), up to a per-episode exploration budget. We then design an algorithm that uses a UCB reinforcement learning policy for exploration, but overrides it as needed to ensure safety with high probability. We experimentally validate our results on a sepsis treatment task, demonstrating that our algorithm can learn while ensuring good performance compared to the baseline policy for every patient.