Abstraction and Generalisation in Reinforcement Learning
Abstraction and Generalisation in Reinforcement Learning
批准号:
2281998
负责人:
金额:
$0.0万
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
未结题
起止时间:
2019 至 --
中文摘要
我正在研究如何使用抽象状态空间来表示任务,以及如何从经验中推断这些状态空间。如果你要向某人描述如何在咖啡馆点咖啡,你可以这样说:“走进门,走到柜台前,看看菜单,决定你想要什么,告诉咖啡师你的订单,然后等着你的咖啡端上来。”在强化学习环境中,这组指令定义了状态之间的一系列转换——也就是说,从咖啡馆外的状态,到咖啡馆内的状态,再到柜台前的状态,等等。然而,这个状态空间与你在世界中采取的行动粒度不同——当你走过门时,你不会在肌肉抽搐的层面上考虑你所采取的每一个动作,甚至不会在摆动腿的层面上考虑你所采取的每一个动作。强化学习代理受限于其环境中定义的最精细的动作,在这个例子中类比为肌肉抽搐。状态抽象是agent发现前面描述的高级状态的过程,通过将其粒度操作分层地组合成高阶技能。我目前正在使用以后继表示为中心的强化学习方法——一种将状态占用和奖励结合起来的表示值方法。这让我们能够探究状态空间的结构,而不会因为观察价值结构而产生奖励的混乱。我目前正在使用对称压缩和时间抽象来构建抽象状态空间
英文摘要
I am studying how tasks can be represented using abstract state spaces, and how these state spaces can be inferred from experience. If you were to describe to someone how to order a coffee at a cafe, you might say something like "walk through the door, go to the counter, look at the menu and decide what you want, tell the barista your order and then wait for your coffee to be served". In a reinforcement learning setting, that set of instructions defines a set of transitions between states - that is, the state of being outside the cafe, to the state of being inside the cafe, to the state of being at the counter, and so on. This state space however is not at the same granularity as the actions you take in the world - when you walk though the door, you are not considering every single action you take at the level of muscle twitches, or even at the level of swinging your legs. Reinforcement learning agents are constrained to the most granular action defined in their environment, which by analogy would be muscle twitches in this instance. State abstraction is the process by which an agent would discover the high-level states described before, by composing its granular actions hierarchically into high-order skills. I am currently using reinforcement learning methods centred around the successor representation - a way of representation values as the conjunction of state occupancies and reward. This allows us to probe the structure of the state spaces without the confound of reward we would have by looking at the value structure. I am currently building abstract state spaces using symmetric compression, and through temporal abstraction
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金