Learning to fly by combining reinforcement learning with behavioural cloning

Learning to fly by combining reinforcement learning with behavioural cloning
复制标题

通过强化学习与行为克隆相结合来学习飞行

DOI:
--
复制
发表时间:
2004
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
C. Sammut
C. Sammut
中科院分区:
--
文献类型:
--
作者:
E. Morales;C. Sammut

文献摘要

被引文献

相似文献

强化学习处理学习最优或接近最优的策略,同时与环境交互。由于搜索空间较大,现有的强化学习方法很难解决具有多个连续变量的应用领域问题。在本文中,我们使用关系表示来定义强大的抽象,允许我们整合领域知识并在其他类似问题中重用以前学习的策略。我们还描述了如何使用行为克隆方法结合探索阶段从人类踪迹中学习有用的行动。由于同一抽象状态可能会导致几个相互冲突的行为,因此强化学习被用来在这个缩减的空间上学习最优策略。实验表明,使用关系表示的行为克隆和强化学习的组合如何强大到足以学习如何在空间中的不同点和不同的湍流条件下驾驶飞机。
Reinforcement learning deals with learning optimal or near optimal policies while interacting with the environment. Application domains with many continuous variables are difficult to solve with existing reinforcement learning methods due to the large search space. In this paper, we use a relational representation to define powerful abstractions that allow us to incorporate domain knowledge and re-use previously learned policies in other similar problems. We also describe how to learn useful actions from human traces using a behavioural cloning approach combined with an exploration phase. Since several conflicting actions may be induced for the same abstract state, reinforcement learning is used to learn an optimal policy over this reduced space. It is shown experimentally how a combination of behavioural cloning and reinforcement learning using a relational representation is powerful enough to learn how to fly an aircraft through different points in space and different turbulence conditions.