Data-Efficient Policy Evaluation Through Behavior Policy Search

Data-Efficient Policy Evaluation Through Behavior Policy Search
复制标题

DOI:
--
复制
发表时间:
2017-06
期刊:
--
影响因子:
--
通讯作者:
Josiah P. Hanna;P. Thomas;P. Stone;S. Niekum
Josiah P. Hanna;P. Thomas;P. Stone;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Josiah P. Hanna;P. Thomas;P. Stone;S. Niekum

文献摘要

被引文献

相似文献

我们考虑的任务是评估马尔可夫决策过程(MDP)的政策。评估策略的标准无偏技术是部署策略并观察其性能。我们发现,从部署一个不同的政策,通常被称为行为政策,收集的数据,可以用来产生无偏估计与较低的均方误差比这个标准的技术。我们推导出一个最优行为策略的解析表达式-最小化所得到的估计的均方误差的行为策略。因为这个表达式依赖于在实践中未知的条款,我们提出了一个新的政策评估子问题,行为政策搜索:寻找一个行为政策,减少均方误差。我们提出了一种行为政策搜索算法,并实证证明了其在降低政策绩效估计均方误差方面的有效性。
We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its performance. We show that the data collected from deploying a different policy, commonly called the behavior policy, can be used to produce unbiased estimates with lower mean squared error than this standard technique. We derive an analytic expression for the optimal behavior policy --- the behavior policy that minimizes the mean squared error of the resulting estimates. Because this expression depends on terms that are unknown in practice, we propose a novel policy evaluation sub-problem, behavior policy search: searching for a behavior policy that reduces mean squared error. We present a behavior policy search algorithm and empirically demonstrate its effectiveness in lowering the mean squared error of policy performance estimates.