Learning Options with Interest Functions

Learning Options with Interest Functions
复制标题

具有兴趣功能的学习选项

DOI:
--
复制
发表时间:
2019
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Doina Precup
Doina Precup
中科院分区:
--
文献类型:
--
作者:
Khimya Khetarpal;Doina Precup

文献摘要

被引文献

相似文献

学习时态抽象是任务的部分解决方案,可以重用于解决其他任务是一个成分,可以帮助代理计划和学习效率。在这项工作中,我们在选项框架中解决这个问题。我们的目标是自主学习的选项,专门在不同的状态空间区域提出了一个概念的兴趣函数,它概括了启动集的选项框架函数逼近。我们建立在期权批评的框架,推导出政策梯度定理的利益函数,导致一个新的利益期权批评架构。
Learning temporal abstractions which are partial solutions to a task and could be reused for solving other tasks is an ingredient that can help agents to plan and learn efficiently. In this work, we tackle this problem in the options framework. We aim to autonomously learn options which are specialized in different state space regions by proposing a notion of interest functions, which generalizes initiation sets from the options framework for function approximation. We build on the option-critic framework to derive policy gradient theorems for interest functions, leading to a new interest-option-critic architecture.