Empirical Policy Evaluation With Supergraphs

Empirical Policy Evaluation With Supergraphs
复制标题

DOI:
10.1109/jsait.2021.3073257
复制
发表时间:
2020-02
期刊:
IEEE Journal on Selected Areas in Information Theory
影响因子:
--
通讯作者:
Daniel Vial;V. Subramanian
Daniel Vial;V. Subramanian
中科院分区:
其他
文献类型:
--
作者:
Daniel Vial;V. Subramanian

文献摘要

相似文献

我们为强化学习中的策略评估问题设计了算法,假设可以访问模拟器和称为超图的某些辅助信息。我们的算法从高成本状态向后探索,以找到高价值状态,这与从所有状态向前工作的方法形成对比。虽然已有几篇论文从经验上证明了反向探索的有效性,但我们进行了严格的分析,结果表明,我们的算法可以将平均样本复杂度从$O(S\log S)$降低到$O(\log S)$。在分析方面,我们采用了网络科学文献中的工具,为强化学习问题提供了一种新的方法。
We devise algorithms for the policy evaluation problem in reinforcement learning, assuming access to a simulator and certain side information called the supergraph. Our algorithms explore backward from high-cost states to find high-value ones, in contrast to approaches that work forward from all states. While several papers have demonstrated the utility of backward exploration empirically, we conduct rigorous analyses which show that our algorithms can reduce average-case sample complexity from $O(S \log S)$ to as low as $O(\log S)$ . Analytically, we adapt tools from the network science literature to provide a new methodology for reinforcement learning problems.