New Algorithms for Markov Decision Processes and Reinforcement Learning
New Algorithms for Markov Decision Processes and Reinforcement Learning
批准号:
2208163
负责人:
Lexing Ying
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2025-08-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Markov decision processes and reinforcement learning have had significant recent success in applications, ranging from outperforming humans in Atari games to AlphaFold overshadowing competing methods in predicting protein folding. This success results from several fundamental developments, including deep neural networks providing a powerful mechanism for representing high dimensional functions, unprecedented computing power provided by graphical processing units and tensor processing units, and the development of novel algorithms for both prediction and control. However, there are still many challenges in applying these recent techniques to mission-critical applications in health, social and economic planning, and defense. This project aims to develop and analyze novel algorithms for Markov decision processes and reinforcement learning with the intention of making these approaches more broadly applicable. Educational impacts include postdoctoral and graduate student training, as well as undergraduate course development centered around machine learning. This project involves the development of a unified framework for Markov decision processes based on linear programming, where the primal, dual, and primal-dual problems are studied for both the regularized and non-regularized cases. Existing algorithms based on Markov decision processes will then be connected to this unified framework. For the tabular setting, a quasi-Newton type policy gradient algorithm will be developed for general entropic regularizers. For the primal-dual problem, a rapidly converging gradient ascent descent algorithm based on a strictly convexified formulation with a non-standard preconditioning metric will be developed. The nonlinear approximation setting will be addressed by variational actor-critic algorithms that are stable and converge at least to a local minimum. Finally, to address the double sampling issue, new algorithms based on the borrowing-from-the-future idea will be developed to significantly reduce the bias.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep Learning for Inverse Problems
-
批准号:2011699
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2020
-
负责人:Lexing Ying
-
依托单位:
Tensor Network Computation: Representations, Algebra, and Applications
-
批准号:1818449
-
项目类别:Continuing Grant
-
资助金额:$18.0万
-
财政年份:2018
-
负责人:Lexing Ying
-
依托单位:
Effective Preconditioners for High Frequency Wave Equations
-
批准号:1521830
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:2015
-
负责人:Lexing Ying
-
依托单位:
CDI-Type I: Collaborative Research: High-Dimensional Phase-Space Subdivisions for Seismic Imaging
-
批准号:1327658
-
项目类别:Standard Grant
-
资助金额:$6.79万
-
财政年份:2013
-
负责人:Lexing Ying
-
依托单位:
CAREER: Fast Algorithms for Oscillatory Integrals
-
批准号:1328230
-
项目类别:Standard Grant
-
资助金额:$21.66万
-
财政年份:2013
-
负责人:Lexing Ying
-
依托单位:
CDI-Type I: Collaborative Research:High-dimensional phase-space subdivisions for seismic imaging
-
批准号:1027952
-
项目类别:Standard Grant
-
资助金额:$9.19万
-
财政年份:2010
-
负责人:Lexing Ying
-
依托单位:
CAREER: Fast Algorithms for Oscillatory Integrals
-
批准号:0846501
-
项目类别:Standard Grant
-
资助金额:$41.5万
-
财政年份:2009
-
负责人:Lexing Ying
-
依托单位:
Collaborative Research: Wave Computations in Phase-Space
-
批准号:0708014
-
项目类别:Standard Grant
-
资助金额:$15.47万
-
财政年份:2007
-
负责人:Lexing Ying
-
依托单位:
海外基金