Improved analysis of policy gradient methods in reinforcement learning.
Improved analysis of policy gradient methods in reinforcement learning.
批准号:
2602524
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
强化学习是机器学习的一个重要分支,旨在解决环境中的序贯决策问题。它具有广泛的应用,包括自动驾驶、机器人、推荐系统和医疗保健。在其中一些应用中,错误决策的代价可能是巨大的。尤其是在自动驾驶等应用中,人类的生命危在旦夕。因此,我们必须了解这些方法是如何工作的,以及它们是否真的以预期的方式工作,这一点至关重要。然而,实践中使用的方法往往只被很少了解。目前描述这些方法的理论无法解释强化学习在实践中取得的巨大成功。该项目的目的是为被称为策略梯度方法的方法提供改进的理论保证,这些方法构成了强化学习的大部分实际实现的基础。这些方法特别适用于实际中经常遇到的大规模问题,具体地说,这类算法的理论具有收敛界的形式。也就是说,算法的目标是输出问题的最优解。我们感兴趣的是了解算法输出接近此最优解的东西的速度有多快,其中贴近度的概念在数学上是精确的。改进分析的目的在这里解释为算法比以前证明的收敛速度更快。最近,在一种称为策略镜像下降的新视角下,研究了特定环境下的一种特定类型的策略梯度方法。这到底意味着什么并不是太重要,除了镜面下降是一个来自优化理论的概念,在那个背景下已经进行了大量的研究。因此,分析的工具和方法可以从最优化理论转换到这种强化学习框架。这可以被用来实现改进的收敛保证,这是我们在这个项目中使用的方法之一。这个项目是StatML CDT的一部分,该CDT是伦敦帝国理工学院和牛津大学的联合CDT。它属于EPSRC统计和应用概率研究领域。特别是,尽管这个项目与优化密切相关,但它仍然具有很强的统计学性质。这是因为为了解决强化学习的决策问题,我们感兴趣的是利用数据本身具有一定的随机性。
英文摘要
Reinforcement learning is a popular branch of machine learning that aims to solve a sequential decision-making problem in an environment. This has a wide variety of applications including autonomous driving, robotics, recommendation systems and healthcare. In some of these applications, the cost of a wrong decision could be dramatic. In particular for applications like autonomous driving, the lives of human beings are at stake. As such, it is of crucial importance that we understand how the methods work and whether they really do work in the way that was intended.However, the methods that are used in practice are often only poorly understood. The theory describing these methods is currently unable to explain the huge successes that reinforcement learning has enjoyed in practice. The aim of this project is to provide improved theoretical guarantees for methods known as policy gradient methods that form the basis for much of the practical implementations of reinforcement learning. These methods are particularly used for large-scale problems that are often faced in practice.Specifically, theory on algorithms of this type takes the form of convergence bounds. That is, the algorithm is aiming to output a solution to the problem that is optimal. We are interested in understanding how quickly the algorithm outputs something close to this optimal solution, where the notion of closeness is mathematically precise. The aim of improved analyses translates here into saying that an algorithm converges faster than what was previously proven.Recently, a particular type of a policy-gradient method in a specific setting has been studied under a new perspective known as policy mirror descent. What exactly this means is not too important except that mirror descent is a concept from optimisation theory that has been heavily studied in that setting. As such, tools and methods of analysis may be translated from optimisation theory to this reinforcement learning framework. This can be exploited to achieve improved convergence guarantees, which is one of the avenues that we are using in this project.This project is part of the StatML CDT, which is a joint CDT between Imperial College London and the university of Oxford. It falls within the EPSRC statistics and applied probability research area. In particular, though this project is heavily linked to optimisation, it remains very statistical in nature. This is because we are interested in using data that inherently has some randomness to it in order to solve the decision-making problem of reinforcement learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:USHARANI HAREESH GOVINDARA JAN
-
依托单位:
利用全基因组关联分析和QTL-seq发掘花生白绢病抗性分子标记
-
批准号:31971981
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:晏立英
-
依托单位:
基于SERS纳米标签和光子晶体的单细胞Western Blot定量分析技术研究
-
批准号:31900571
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:刘兵
-
依托单位:
利用多个实验群体解析猪保幼带形成及其自然消褪的遗传机制
-
批准号:31972542
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2019
-
负责人:郭源梅
-
依托单位:
基于Meta-analysis的新疆棉花灌水增产模型研究
-
批准号:41601604
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:赵爱琴
-
依托单位:
基于个体分析的投影式非线性非负张量分解在高维非结构化数据模式分析中的研究
-
批准号:61502059
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2015
-
负责人:刘昶
-
依托单位:
多目标诉求下我国交通节能减排市场导向的政策组合选择研究
-
批准号:71473155
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2014
-
负责人:柴建
-
依托单位:
大规模微阵列数据组的meta-analysis方法研究
-
批准号:31100958
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:赵洪雅
-
依托单位:
基于物质流分析的中国石油资源流动过程及碳效应研究
-
批准号:41101116
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2011
-
负责人:刘晓洁
-
依托单位: