Improving Sample Efficiency of Reinforcement Learning
Improving Sample Efficiency of Reinforcement Learning
批准号:
2579743
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
深度强化学习在经验上取得了巨大的成功,是许多人工智能应用的主要使能技术。然而,最近的强化学习算法仍然需要数百万个样本才能获得良好的性能。由于获得环境相互作用通常是昂贵的,并且具有挑战性的环境很少是静态的,因此这抑制了许多实际应用。该项目将研究降低该成本的方法,旨在找到更高效的样本强化学习算法。我们的目标是在现实环境中部署算法,其中代理使用深度网络来表示有关环境的知识。它还可能导致其他系统做出自动化决策的性能得到改善。研究策略该项目将研究提高样品效率的两个主要途径。首先,使用贝叶斯框架从样本中获得额外的信息,我们希望实现改进的探索,从而获得更多信息的样本。其次,使用元学习,我们希望实现泛化,减少学习与其他学习任务相似的任务所需的样本数量。目标和应用我们的目标是开发更多的样本高效强化学习算法,并获得关于探索的新见解。该项目将与微软剑桥研究院合作进行,并将与他们的电脑游戏研究直接相关,特别是在样本昂贵的复杂世界中训练游戏AI。该项目将有更广泛的应用于任何涉及有限数据决策的问题,包括现实世界的应用,如机器人和定价策略。
英文摘要
Deep reinforcement learning has had huge empirical success and is a major enabling technology for many applications of AI. However, recent RL algorithms still require millions of samples to obtain good performance. Since obtaining environment interactions is often costly and since challenging environments are rarely static, this inhibits many practical applications. This project will investigate ways of reducing this cost, aiming to find more sample-efficient RL algorithms. We aim for the algorithms to be deployable in realistic settings, where agents use deep networks to represent knowledge about the environment. It is also likely to lead to improved performance of other systems making automated decisions. Research StrategyThe project will investigate two main avenues for improving sample efficiency. Firstly, using a Bayesian framework to gain additional information from samples, we hope to achieve improved exploration which will in turn lead to more informative samples. Secondly, using meta-learning, we hope to enable generalisation, reducing the amount of samples required to learn a task which is similar to other learned tasks.Objectives and ApplicationsWe aim to develop more sample efficient reinforcement learning algorithms and to gain new insights about exploration. The project will be carried out in collaboration with Microsoft Research Cambridge and will have immediate relevance for their computer games research, particularly for training game AI in complex worlds where samples are expensive. The project will have wider applications for any problem which involves decision making with limited data, including real world applications such as robotics and pricing strategies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金