Safely Bridging Offline and Online Reinforcement Learning

Safely Bridging Offline and Online Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Wanqiao Xu;Kan Xu;Hamsa Bastani;O. Bastani
Wanqiao Xu;Kan Xu;Hamsa Bastani;O. Bastani
中科院分区:
其他
文献类型:
--
作者:
Wanqiao Xu;Kan Xu;Hamsa Bastani;O. Bastani

文献摘要

被引文献

相似文献

在实践中部署强化学习的一个关键挑战是安全地探索。我们提出了一个自然的安全属性-一致优于保守的政策(自适应估计从迄今为止观察到的所有数据),每集的勘探预算。然后,我们设计了一个算法,该算法使用UCB强化学习策略进行探索,但根据需要覆盖它,以确保高概率的安全性。我们通过实验验证了我们在脓毒症治疗任务上的结果,表明我们的算法可以学习,同时确保与每个患者的基线策略相比具有良好的性能。
A key challenge to deploying reinforcement learning in practice is exploring safely. We propose a natural safety property— uniformly outperforming a conservative policy (adaptively estimated from all data observed thus far), up to a per-episode exploration budget. We then design an algorithm that uses a UCB reinforcement learning policy for exploration, but overrides it as needed to ensure safety with high probability. We experimentally validate our results on a sepsis treatment task, demonstrating that our algorithm can learn while ensuring good performance compared to the baseline policy for every patient.