Mechanizing Soundness of Off-Policy Evaluation

Mechanizing Soundness of Off-Policy Evaluation
复制标题

DOI:
10.4230/lipics.itp.2022.32
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Jared Yeager;J. Moss;Michael Norrish;P. Thomas
Jared Yeager;J. Moss;Michael Norrish;P. Thomas
中科院分区:
其他
文献类型:
--
作者:
Jared Yeager;J. Moss;Michael Norrish;P. Thomas

文献摘要

被引文献

相似文献

有一些强化学习的场景——例如,在医学领域——在实施政策之前,我们被迫尽可能地相信政策的改变会带来改善。在这种情况下,我们可以使用策略外评估(OPE)。OPE的基本思想是记录当前政策下的行为历史,然后对提议的新政策的质量进行评估,看看在新政策下这些行为会是什么样子。由于我们正在评估策略而没有实际使用它,因此我们有OPE的“off-policy”。将集中不等式应用于估计,我们得到了新政策预期质量的置信区间。如果置信区间高于当前政策的置信区间,我们就可以很有信心地改变政策,不会造成伤害。我们在这里着重于这种方法的数学,通过机械化的健全的非政策评估。机械化的一个自然副作用是既澄清了所有结果的数学假设和先决条件,又进一步发展了HOL4的经过验证的统计数学库,包括浓度不等式。更重要的是,OPE方法依赖于重要抽样,我们用测量理论的方法证明了它的合理性。事实上,我们推广了标准结果,在包含离散和连续概率分布的情况下展示了它。
There are reinforcement learning scenarios – e.g., in medicine – where we are compelled to be as confident as possible that a policy change will result in an improvement before implementing it. In such scenarios, we can employ off-policy evaluation (OPE). The basic idea of OPE is to record histories of behaviors under the current policy, and then develop an estimate of the quality of a proposed new policy, seeing what the behavior would have been under the new policy. As we are evaluating the policy without actually using it, we have the “off-policy” of OPE. Applying a concentration inequality to the estimate, we derive a confidence interval for the expected quality of the new policy. If the confidence interval lies above that of the current policy, we can change policies with high confidence that we will do no harm. We focus here on the mathematics of this method, by mechanizing the soundness of off-policy evaluation. A natural side effect of the mechanization is both to clarify all the result’s mathematical assumptions and preconditions, and to further develop HOL4’s library of verified statistical mathematics, including concentration inequalities. Of more significance, the OPE method relies on importance sampling, whose soundness we prove using a measure-theoretic approach. In fact, we generalize the standard result, showing it for contexts comprising both discrete and continuous probability distributions.