Towards Safe Policy Improvement for Non-Stationary MDPs

Towards Safe Policy Improvement for Non-Stationary MDPs
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yash Chandak;Scott M. Jordan;Georgios Theocharous;Martha White;P. Thomas
Yash Chandak;Scott M. Jordan;Georgios Theocharous;Martha White;P. Thomas
中科院分区:
其他
文献类型:
--
作者:
Yash Chandak;Scott M. Jordan;Georgios Theocharous;Martha White;P. Thomas

文献摘要

被引文献

相似文献

许多现实世界中的序贯决策问题涉及具有金融风险和人类生命风险的关键系统。虽然在过去的几个作品已经提出了方法,是安全的部署,他们假设的根本问题是固定的。然而,许多现实世界的问题表现出非平稳性,当风险很高时,与错误的平稳性假设相关的成本可能是不可接受的。我们采取的第一步,以确保安全性,高信心,平稳变化的非平稳决策问题。我们提出的方法扩展了一种安全的算法,称为Seldonian算法,通过无模型强化学习与时间序列分析的合成。安全性是确保使用顺序假设检验的政策的预测性能,并获得置信区间使用野生自助。
Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that are safe for deployment, they assume that the underlying problem is stationary. However, many real-world problems of interest exhibit non-stationarity, and when stakes are high, the cost associated with a false stationarity assumption may be unacceptable. We take the first steps towards ensuring safety, with high confidence, for smoothly-varying non-stationary decision problems. Our proposed method extends a type of safe algorithm, called a Seldonian algorithm, through a synthesis of model-free reinforcement learning with time-series analysis. Safety is ensured using sequential hypothesis testing of a policy's forecasted performance, and confidence intervals are obtained using wild bootstrap.