课题基金 / 基金详情

Hierarchical reinforcement learning in large-scale domains

Hierarchical reinforcement learning in large-scale domains
大规模领域的分层强化学习
批准号:
2120604
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
时间抽象对于智能系统是有价值的。在多个时间尺度上表示知识提供了一种划分状态空间的方法,这可以加速学习并允许将行为转移到不同的任务。人类不断地使用时间上延伸的行动来计划和行动,将任何特定的任务分解为一系列突出的路点或子目标。分层强化学习依赖于一组理论上合理的方法,用于使用时间扩展动作进行学习和规划[1],[2],[3]。尽管我们有丰富的表达框架,利用一个给定的层次结构的行动,一个主要的问题,仍然是我们如何可以自主地发现一个给定的域的层次结构。这个问题被称为技能发现。存在许多很好的技能发现方法,其中一些基于图论,一些基于挖掘强化学习代理经验的轨迹,还有一些基于梯度的端到端优化。然而,目前的方法并不能立即与所有类型的问题兼容,并且还没有被证明可以很好地扩展。魔方是一个标志性的难题,对于没有先验知识的人来说很难解决。目前还没有解决方案使用从任意置乱状态开始的强化学习。这个问题的一个明显的元素是需要同时满足竞争目标,即通过正确放置立方体的不同部分。另一个困难的关键来源是由于不可序列化的子目标的属性:使用一系列子目标来获得解决方案,在达到进一步的子目标之前,必须暂时违反一些先前的子目标。虽然众所周知,魔方可以在20步或更少的时间内从其43个quintillion状态中的任何一个状态中解决,但可以解决立方体的“立体主义者”通常通过采用各种宏运算符来使用更多的移动。这些宏操作符使状态的一部分对其效果保持不变,这允许立体派在求解的每个阶段仅操作魔方的某些部分。这项研究将集中在如何强化学习代理可以学习的层次策略的魔方的问题。所进行的初步工作已经确定了状态空间的一个关键属性。未来可能的方向可以解决从直接经验中发现宏操作符,开发限制初始集的方法,并利用问题的对称性。需要仔细考虑设计有效的功能近似方法,无论是在最高控制级别还是在临时扩展的行动中。除了魔方之外,还有许多排列谜题也可以通过本研究创建的方法来解决。更一般地说,组合优化问题在科学和工程中广泛存在,并且越来越多地使用强化学习来解决[5]。目的是将这项研究所产生的方法纳入这一更广泛的工作中。[1]帕尔河,Russell,S. 1998.强化学习与机器层次结构。神经信息处理系统的进展:第10届会议论文集,丹佛。麻省理工学院出版社. [2]萨顿河美国,普雷卡普,D.,和Singh,S. 1999.在MDP和Semi-MDP之间:强化学习中的时间抽象框架。人工智能,112,第181 -211页。[3]Dietterich,T. G. 2000.分层强化学习与MAXQ值函数分解。人工智能研究杂志,13,pp。227-303. [4]科尔夫河1985. Macro-operators:一种弱学习方法。人工智能,35,pp。35-77. [5]燕军湖,亨通,K.,Ketian,Y.,Shuyu,Y.,Xiaolin,L. 2018.零折叠:疏水极性模型中的蛋白质零折叠。神经信息处理系统的进展
英文摘要
Temporal abstraction is valuable for an intelligent system. Representing knowledge over multiple timescales provides a means of partitioning state space, which can accelerate learning and allow behaviour to be transferred to different tasks. Humans constantly plan and behave using temporally extended actions, breaking any particular task down into a sequence of salient waypoints, or subgoals. Hierarchical reinforcement learning rests upon a set of theoretically sound approaches for learning and planning using temporally extended actions [1], [2], [3]. Despite us having richly expressive frameworks for utilising a given hierarchy of action, a major problem that remains is how we may autonomously discover the hierarchical structure of a given domain. This problem is known as skill discovery. There exist many good approaches to skill discovery, with some based on graph theory, some on mining the trajectories of a reinforcement learning agent's experience, and others on gradient based, end-to-end optimisation. However, the current methods are not immediately compatible with all types of problems, and have not been demonstrated to scale well.Rubik's cube is an iconic puzzle that has a reputation for being difficult to solve for someone without prior knowledge. There are currently no solutions that use reinforcement learning starting from an arbitrary scrambled state. An obvious element of the problem is the need to simultaneously satisfy competing objectives, i.e. by correctly placing the different pieces of the cube. Another key source of difficulty is due to the property of non-serialisable subgoals: using a sequence of subgoals to arrive at the solution, some previous subgoals must be temporarily violated before reaching further ones. Whilst it is known that Rubik's cube can be solved in 20 moves or less from any of its 43 quintillion states, 'cubists' who can solve the cube typically use many more moves by employing a variety of macro operators [4]. These macro operators leave part of a state invariant to their effects, which allows cubists to manipulate only certain parts of Rubik's cube during each stage of their solve. This research will focus on the question of how a reinforcement learning agent may learn a hierarchical policy for Rubik's cube. Preliminary work undertaken has identified a key property of the state space. Possible future directions could address the discovery of macro operators from direct experience, develop ways to restrict initiation sets, and utilise symmetries of the problem. Careful consideration will be needed to design effective methods of function approximation, both at the top-level of control and also for the temporarily extended actions. Beyond the Rubik's cube there are many permutation puzzles that can also be solved through methods this research will create. More generally, combinatorial optimisation problems are widespread throughout science and engineering, and are increasingly being addressed using reinforcement learning [5]. The aim is to incorporate methods arising from this research into this wider body of work. [1] Parr, R., and Russell, S. 1998. Reinforcement learning with hierarchies of machines. In Advances in Neural Information Processing Systems: Proceedings of the 10th Conference, Denver. Cambridge, MA: MIT Press. [2] Sutton, R. S., Precup, D., and Singh, S. 1999. Between MDPs and Semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112, pp.181-211.[3] Dietterich, T. G. 2000. Hierarchical reinforcement learning with the MAXQ value function decomposition. Journal of Artificial Intelligence Research, 13, pp. 227-303.[4] Korf, R. 1985. Macro-operators: A weak method for learning. Artificial Intelligence, 35, pp. 35-77.[5] Yanjun, L., Hengtong, K., Ketian, Y., Shuyu, Y., and Xiaolin, L. 2018. FoldingZero: Protein Folding from Scratch in Hydrophobic-Polar Model. In Advances in Neural Information Processing Systems
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
海桑属杂种区强化(Reinforcement)的检验与遗传基础研究
  • 批准号:
    30800060
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2008
  • 负责人:
    周仁超
  • 依托单位: