课题基金 / 基金详情

Improved analysis of policy gradient methods in reinforcement learning.

Improved analysis of policy gradient methods in reinforcement learning.
强化学习中策略梯度方法的改进分析。
批准号:
2602524
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Reinforcement learning is a popular branch of machine learning that aims to solve a sequential decision-making problem in an environment. This has a wide variety of applications including autonomous driving, robotics, recommendation systems and healthcare. In some of these applications, the cost of a wrong decision could be dramatic. In particular for applications like autonomous driving, the lives of human beings are at stake. As such, it is of crucial importance that we understand how the methods work and whether they really do work in the way that was intended.However, the methods that are used in practice are often only poorly understood. The theory describing these methods is currently unable to explain the huge successes that reinforcement learning has enjoyed in practice. The aim of this project is to provide improved theoretical guarantees for methods known as policy gradient methods that form the basis for much of the practical implementations of reinforcement learning. These methods are particularly used for large-scale problems that are often faced in practice.Specifically, theory on algorithms of this type takes the form of convergence bounds. That is, the algorithm is aiming to output a solution to the problem that is optimal. We are interested in understanding how quickly the algorithm outputs something close to this optimal solution, where the notion of closeness is mathematically precise. The aim of improved analyses translates here into saying that an algorithm converges faster than what was previously proven.Recently, a particular type of a policy-gradient method in a specific setting has been studied under a new perspective known as policy mirror descent. What exactly this means is not too important except that mirror descent is a concept from optimisation theory that has been heavily studied in that setting. As such, tools and methods of analysis may be translated from optimisation theory to this reinforcement learning framework. This can be exploited to achieve improved convergence guarantees, which is one of the avenues that we are using in this project.This project is part of the StatML CDT, which is a joint CDT between Imperial College London and the university of Oxford. It falls within the EPSRC statistics and applied probability research area. In particular, though this project is heavily linked to optimisation, it remains very statistical in nature. This is because we are interested in using data that inherently has some randomness to it in order to solve the decision-making problem of reinforcement learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    USHARANI HAREESH GOVINDARA JAN
  • 依托单位:
利用全基因组关联分析和QTL-seq发掘花生白绢病抗性分子标记
基于SERS纳米标签和光子晶体的单细胞Western Blot定量分析技术研究
  • 批准号:
    31900571
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2019
  • 负责人:
    刘兵
  • 依托单位: