课题基金 / 基金详情

Scaling Unsupervised Environment Design

Scaling Unsupervised Environment Design
扩展无监督环境设计
批准号:
2888076
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
强化学习(RL)是机器学习的一个子领域,其中智能体(例如自动驾驶汽车)通过在环境(例如真实道路/道路模拟)中的行为进行学习。尽管在解决复杂的视频游戏(雅达利、围棋、星际争霸)方面取得了很大的进步,但它还没有成功地应用于许多现实世界的问题。造成这种情况的根本原因是RL代理无法将其概括为不可见的场景。具体地说,由于模拟不可避免的不准确性,在模拟中训练的RL代理在部署在现实世界中时不能很好地传输(注意,由于需要大量的训练数据和潜在的危险,在现实世界中训练代理通常是不切实际的)。最近的开创性工作已经证明了推广的显著经验益处,方法是培训一名教师,该教师学习为代理提出高质量的场景(例如道路布局)进行培训,反映了监督学习的结果,这些结果表明了数据质量在泛化中的重要性。这项工作的一个限制是,教师必须从稀疏和噪声信号中学习,导致样本效率低,需要大量计算资源,这意味着它只成功地应用于非常简单的问题。为了减少信号噪音,我提出了一些方法,鼓励教师在学习的潜在空间中使用近似惊讶、容易辨别和距离的度量来维持一组多样化的场景。此外,本文还提出了一种新颖的数据增强方法,将场景分解为一组子场景,以最小的计算代价扩展训练数据。最后,目前最先进的方法通过应用随机扰动来培训教师。我建议一种有针对性的干扰方法,方法是不断逼近代理的遗憾(它在任务中做得有多好与最佳代理应该做得有多好之间的差异),并在这一点最低的地方应用干扰。所有这些技术都旨在提高整个过程的效率,减少所需的资源,并允许这项强大的技术向更复杂的领域开放,使自动驾驶等现实世界的应用受益。应该注意的是,虽然我使用自动驾驶作为一个运行示例,但正在开发的方法将适用于任何RL问题,并将在不同的环境中进行评估。该项目属于EPSRC人工智能技术研究领域。
英文摘要
Reinforcement learning (RL) is a subfield of machine learning where an agent (e.g. an autonomous vehicle) learns from acting in an environment (e.g. a real road / simulation of a road). Despite making great progress in solving complex video games (Atari, Go, StarCraft), it has not yet been successfully applied to many real world problems. The root cause of this is the inability of RL agents to generalise to unseen scenarios. Specifically, an RL agent trained in a simulation doesn't transfer well when deployed in the real world, due to the inevitable inaccuracies of simulation (note that, due to the large volume of training data needed and potential dangers, it is often impractical to train an agent in the real world).Recent pioneering work has demonstrated significant empirical benefits to generalisation by training a teacher that learns to propose high-quality scenarios (e.g. road layouts) for the agent to train on, mirroring results from supervised learning that have shown the importance of data quality in generalisation. A limitation to this work is that the teacher has to learn from a sparse and noisy signal, resulting in low sample efficiency and necessitating large computational resources, meaning it has only been successfully applied to very simple problems. To reduce signal noise, I have proposed methods encouraging the teacher to maintain a diverse set of scenarios using metrics for approximated surprise, ease of discrimination and distance in a learned latent space. Furthermore, I propose a novel data augmentation method, whereby scenarios are decomposed into a set of 'sub-scenarios', expanding the training data with minimal computational cost. Finally, the current state of the art method trains the teacher by applying random perturbations. I suggest a method for targeted perturbations by constantly approximating the agent's regret (the difference between how well it did at the task and how well an optimal agent would have done) and applying perturbations where this is lowest. All these techniques aim to improve the efficiency of the overall process, reducing the resources needed and allowing this powerful technique to be opened up to more complex domains, benefiting real world applications like autonomous driving. It should be noted that, while I use autonomous driving as a running example, the methods being developed will be generalisable to any RL problem and will be evaluated over a diverse range of environments.This project falls within the EPSRC Artificial intelligence technologies research area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金