Empirical Policy Evaluation With Supergraphs
Empirical Policy Evaluation With Supergraphs
复制标题
DOI:
10.1109/jsait.2021.3073257
复制
发表时间:
2020-02
期刊:
影响因子:
--
通讯作者:
Daniel Vial;V. Subramanian
中科院分区:
文献类型:
--
作者:
Daniel Vial;V. Subramanian
We devise algorithms for the policy evaluation problem in reinforcement learning, assuming access to a simulator and certain side information called the supergraph. Our algorithms explore backward from high-cost states to find high-value ones, in contrast to approaches that work forward from all states. While several papers have demonstrated the utility of backward exploration empirically, we conduct rigorous analyses which show that our algorithms can reduce average-case sample complexity from $O(S \log S)$ to as low as $O(\log S)$ . Analytically, we adapt tools from the network science literature to provide a new methodology for reinforcement learning problems.