Combining Parametric and Nonparametric Models for Off-Policy Evaluation

Combining Parametric and Nonparametric Models for Off-Policy Evaluation
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Omer Gottesman;Yao Liu;Scott Sussex;E. Brunskill;F. Doshi-Velez
Omer Gottesman;Yao Liu;Scott Sussex;E. Brunskill;F. Doshi-Velez
中科院分区:
其他
文献类型:
--
作者:
Omer Gottesman;Yao Liu;Scott Sussex;E. Brunskill;F. Doshi-Velez

文献摘要

相似文献

我们考虑了一种基于模型的方法来执行强化学习中的批量离线评估。我们的方法采用混合专家的方法来结合联合收割机的参数和非参数模型的环境,使最终的价值估计具有最小的预期误差。我们这样做,首先估计每个模型的局部精度,然后使用一个规划器来选择在每个时间步使用哪个模型,以最小化返回误差估计沿着整个轨迹。在各种领域中,我们基于混合的方法优于单独的单个模型以及最先进的基于重要性采样的估计器。
We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-parametric models of the environment such that the final value estimate has the least expected error. We do so by first estimating the local accuracy of each model and then using a planner to select which model to use at every time step as to minimize the return error estimate along entire trajectories. Across a variety of domains, our mixture-based approach outperforms the individual models alone as well as state-of-the-art importance sampling-based estimators.