Redesigning the ways we learn and express policies for autonomous agents, to increasing robustness, resolve simulation-to-real transfer, and enable sh
Redesigning the ways we learn and express policies for autonomous agents, to increasing robustness, resolve simulation-to-real transfer, and enable sh
批准号:
2117714
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
这个项目属于EPSRC机器人和人工智能技术研究领域(工程主题)的福尔斯。当我开始我的博士之旅时,有两个广泛的想法,我相信它们是相互联系的,我正在努力为之做出贡献并进一步发展。首先是开发执行分层强化学习的方法。强化学习(RL)是一种机器学习/人工智能方法,自世纪中期以来一直在研究,但在过去的四年左右才取得了显著的成功,也许最值得注意的是在2016年,一个基于RL的代理在世界各地报道的锦标赛中击败了世界围棋冠军李世石。在牛津机器人研究所,我们非常有兴趣将RL应用于物理机器人控制,以及尚未看到真正有影响力的成功领域。强化学习的一个局限性是,它的训练速度非常慢,数据密集型,并且导致非常脆弱的控制策略,这些策略往往在除了他们已经学会解决的确切问题之外的任何事情上都表现不佳。这些限制阻碍了这种令人兴奋的方法在机器人领域的应用,我的目标是通过构建一种执行分层RL的方法来帮助解决这个问题。这将为RL引入一种新的结构,其中将发生策略和学习以及多个时间尺度或“时间抽象”级别。这将借鉴更成熟和成功的控制理论领域的思想,其中分层控制长期以来一直被应用,但将这些思想带入我们新的基于学习的范式。其次,我希望通过学习在不同层次的时间抽象上进行控制,我就能发明一种机器人-一种以编码最佳计算的机器人运动为中心的方法,以这种方式,它们可以在不同的场景中自动排列和重复使用,不同的机器人我的一位主管,Ioannis Havoutis,参与了一个名为MEMMO(运动记忆)的欧洲项目,该项目正在寻求一种方法,根据预先存在的优化计算和保存(在“运动记忆”中)的运动序列,为具有手臂和腿的复杂机器人执行运动生成。至关重要的是,所寻求的方法应该允许这些序列在不同的机器人平台上设计或执行,然后使用MEMMO系统进行运动。这就需要解决这样一个问题,即以某种方式对来自运动的有用信息进行编码,以便它可以被另一个系统利用。目前的方法没有充分考虑如何根据预期的应用最好地做到这一点,而是利用传统的方法,如自动编码或其他降维方法来“编码”这些运动。当我们在新的情况下利用、改变和重复使用这些动作时,这并没有给我们带来更多的优势。相反,这些方法只是数据压缩。相比之下,这是一个我希望通过开发一种成熟的方法来解决的问题,该方法围绕在多个时间尺度上学习运动和控制,这种方法将从一开始就考虑这类数据的时变性质。
英文摘要
This project falls within the EPSRC Robotics and Artificial Intelligence Technologies research areas (Engineering theme).As I start my PhD journey, there are two broad ideas, which I believe to be linked, which I am working towards contributing towards and developing further. The first is to develop methods to perform hierarchical reinforcement learning. Reinforcement learning (RL) is a machine learning/artificial intelligence methodology which has been researched since the mid-20th century and yet which has only had marked success in the last four or so years, perhaps most notably in 2016 when an RL-based agent beat the world Go champion Lee Sedol in a tournament reported around the world. At the Oxford Robotics Institute, we are very interested in applying RL to physical robotic control, and area where it is yet to see really impactful success. A limitation of RL is that it is immensely slow and data-intensive to train, and results in very brittle control policies that tend to perform badly on anything other than the exact problem they have learnt to solve. These limitations hold back the application of this exciting methodology in the exciting domain of robotics, and my aim is to help tackle this by building a method to perform hierarchical RL. This would introduce a new structure to RL in which policies and learning would occur and multiple timescales, or levels of 'temporal abstraction'. This would draw on ideas from the more established and successful domain of control theory, in which hierarchical control has long been applied, but bring these ideas into our new, learning based paradigm.Secondly, I hope that through work on learning to control at different levels of temporal abstraction, I will be able to devise a robotics-centric approach to encoding optimally computed robot motions in such a way that they can be automatically permuted and reused across different scenarios and different robots. One of my supervisors, Ioannis Havoutis, is involved in a European project called MEMMO (memory of motion), which is seeking to find a way of performing motion generation for complex robots with arms and legs, based on pre-existing optimally computed and saved (in the 'memory of motion') motion sequences. Crucially, the method sought should allow these sequences to have been designed for or executed on different robotic platforms to the one that will then use the MEMMO system for motion. This requires a solution to the problem of encoding in some way the helpful information from the motions such that it can be drawn on by another system. Current approaches do not consider adequately how best to do this in light of the intended application, instead drawing on conventional approaches such as autoencoding or other methods of dimensionality reduction to 'encode' these motions. This gives us no more advantage when it comes to drawing on, altering and reusing these motions in new situations. Rather, these approaches are simply data compression. Contrastingly, this is a problem I hope to tackle through developing a mature methodology around learning motions and control at multiple timescales, an approach which would consider from the very first the time-varying nature of this sort of data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金