Scaling Unsupervised Environment Design
Scaling Unsupervised Environment Design
批准号:
2888076
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
强化学习(RL)是机器学习的一个子领域,其中智能体(例如,自主车辆)通过在环境(例如,真实的道路/道路的模拟)中的动作来学习。尽管在解决复杂的电子游戏(雅达利,围棋,星际争霸)方面取得了很大进展,但它还没有成功地应用于许多真实的世界问题。其根本原因是RL代理无法推广到看不见的场景。具体来说,由于模拟不可避免的不准确性,在模拟中训练的RL代理在部署到真实的世界中时不能很好地转移(注意,由于需要大量的培训数据和潜在的危险,在真实的世界中训练一个智能体通常是不切实际的)。最近的开创性工作已经证明,通过训练一个学会提出高水平建议的教师,质量场景(例如道路布局),用于智能体进行训练,反映监督学习的结果,这些结果显示了数据质量在泛化中的重要性。这项工作的一个局限性是,教师必须从稀疏和嘈杂的信号中学习,导致采样效率低,需要大量的计算资源,这意味着它只能成功地应用于非常简单的问题。为了减少信号噪声,我提出了一些方法,鼓励教师在学习的潜在空间中使用近似惊喜,易于区分和距离的度量来保持各种场景。此外,我提出了一种新的数据增强方法,将场景分解为一组“子场景”,以最小的计算成本扩展训练数据。最后,目前的最先进的方法通过应用随机扰动来训练教师。我提出了一种方法,通过不断逼近代理的遗憾(它在任务中做得有多好和最佳代理会做得有多好之间的差异),并在最低的地方应用扰动来进行有针对性的扰动。所有这些技术都旨在提高整个过程的效率,减少所需的资源,并允许这种强大的技术开放给更复杂的领域,使自动驾驶等真实的应用受益。值得注意的是,虽然我使用自动驾驶作为运行示例,但正在开发的方法将可推广到任何RL问题,并将在各种环境下进行评估。该项目属于EPSRC人工智能技术研究领域。福尔斯
英文摘要
Reinforcement learning (RL) is a subfield of machine learning where an agent (e.g. an autonomous vehicle) learns from acting in an environment (e.g. a real road / simulation of a road). Despite making great progress in solving complex video games (Atari, Go, StarCraft), it has not yet been successfully applied to many real world problems. The root cause of this is the inability of RL agents to generalise to unseen scenarios. Specifically, an RL agent trained in a simulation doesn't transfer well when deployed in the real world, due to the inevitable inaccuracies of simulation (note that, due to the large volume of training data needed and potential dangers, it is often impractical to train an agent in the real world).Recent pioneering work has demonstrated significant empirical benefits to generalisation by training a teacher that learns to propose high-quality scenarios (e.g. road layouts) for the agent to train on, mirroring results from supervised learning that have shown the importance of data quality in generalisation. A limitation to this work is that the teacher has to learn from a sparse and noisy signal, resulting in low sample efficiency and necessitating large computational resources, meaning it has only been successfully applied to very simple problems. To reduce signal noise, I have proposed methods encouraging the teacher to maintain a diverse set of scenarios using metrics for approximated surprise, ease of discrimination and distance in a learned latent space. Furthermore, I propose a novel data augmentation method, whereby scenarios are decomposed into a set of 'sub-scenarios', expanding the training data with minimal computational cost. Finally, the current state of the art method trains the teacher by applying random perturbations. I suggest a method for targeted perturbations by constantly approximating the agent's regret (the difference between how well it did at the task and how well an optimal agent would have done) and applying perturbations where this is lowest. All these techniques aim to improve the efficiency of the overall process, reducing the resources needed and allowing this powerful technique to be opened up to more complex domains, benefiting real world applications like autonomous driving. It should be noted that, while I use autonomous driving as a running example, the methods being developed will be generalisable to any RL problem and will be evaluated over a diverse range of environments.This project falls within the EPSRC Artificial intelligence technologies research area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金