CIF: Small: How Much of Reinforcement Learning is Gradient Descent?
CIF: Small: How Much of Reinforcement Learning is Gradient Descent?
批准号:
2245059
负责人:
Alexander Olshevsky
金额:
$30.12万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2026-05-31
中文摘要
在过去的十年里,强化学习在广泛的应用中取得了显著的成功,从国际象棋和围棋等游戏到芯片设计和航空导航等高级应用。现在有充分的证据表明,强化学习代表着提供下一代自主系统的最有前途的研究方向之一。然而,许多流行的强化学习方法往往无法收敛,这使得强化学习在实践中的使用更像是一门艺术,而不是一门科学。这个项目将探索一种新的方法来分析和设计收敛的强化学习方法,该方法基于最近发现的与梯度下降的联系。这种联系不仅将改进现有算法的分析,还将导致新方法的发展。该项目建立在一个新的概念-梯度分裂的基础上,该概念允许将经典的强化学习方法视为对随机梯度下降更新的修改,该方法继承了梯度下降的许多关键性质。我们将使用这种联系来开发时差学习和Q学习的变体,当给定从马尔可夫决策过程中采样的数据集时,它们将几何地收敛于真值函数的统计最优估计。与神经网络近似相结合,我们的方法将用与底层神经网络宽度的幂成反比的额外误差来逼近真值函数。这些结果将被用来开发一种可证明收敛的神经行动者-批评者方法。我们将开发的新方法不仅将对神经网络在强化学习中的性能提供严格的限制,而且将导致比现有方法更快的训练时间。该奖项反映了NSF的法定使命,并已通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the past decade, reinforcement learning has achieved remarkable success in a wide range of applications, from games such as chess and go to advanced applications such as chip design and aerial navigation. There is now ample evidence that reinforcement learning represents one of the most promising research directions to deliver the next generation of autonomous systems. However, many popular reinforcement-learning methods often fail to converge, making the use of reinforcement learning in practice more an art than a science. This project will explore a novel approach to analyzing and designing convergent reinforcement-learning methods based on a recently discovered connection to gradient descent. This connection will not only improve the analysis of existing algorithms but also lead to the development of new methods.This project builds on a novel concept, gradient splitting, which allows classical reinforcement-learning methods to be viewed as modifications of stochastic-gradient-descent updates, which inherit many key properties of gradient descent. We will use this connection to develop variations of temporal difference learning and Q-learning which, when given a dataset sampled from a Markov decision process, will converge geometrically to the statistically optimal estimate of the true value function. Coupled with neural-network approximation, our methods will approximate the true value function with an additional error that is inversely proportional to a power of the width of the underlying neural network. These results will then be used to develop a provably convergent neural actor-critic method. The new methods we will develop will not only provide rigorous bounds on the performance of neural networks in reinforcement learning but also will result in significantly faster training times compared to existing methods.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/lcsys.2023.3287952
发表时间:
2021-04
期刊:
IEEE Control Systems Letters
影响因子:
3
作者:
[R. Liu;Alexander Olshevsky]
通讯作者:
R. Liu;Alexander Olshevsky
CPS: Medium: Federated Learning for Predicting Electricity Consumption with Mixed Global/Local Models
-
批准号:2317079
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2024
-
负责人:Alexander Olshevsky
-
依托单位:
Computationally Efficient Methods for Control of Epidemics on Networks
-
批准号:2240848
-
项目类别:Standard Grant
-
资助金额:$35.24万
-
财政年份:2023
-
负责人:Alexander Olshevsky
-
依托单位:
Efficiently Distributing Optimization over Large-Scale Networks
-
批准号:1933027
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2019
-
负责人:Alexander Olshevsky
-
依托单位:
CAREER: Algorithms and Fundamental Limitations for Sparse Control
-
批准号:1740451
-
项目类别:Standard Grant
-
资助金额:$24.91万
-
财政年份:2017
-
负责人:Alexander Olshevsky
-
依托单位:
Achieving Consensus Among Autonomous Dynamic Agents using Control Laws that Maintain Performance as Network Size Increases
-
批准号:1740452
-
项目类别:Standard Grant
-
资助金额:$15.21万
-
财政年份:2016
-
负责人:Alexander Olshevsky
-
依托单位:
Achieving Consensus Among Autonomous Dynamic Agents using Control Laws that Maintain Performance as Network Size Increases
-
批准号:1463262
-
项目类别:Standard Grant
-
资助金额:$30.09万
-
财政年份:2015
-
负责人:Alexander Olshevsky
-
依托单位:
CAREER: Algorithms and Fundamental Limitations for Sparse Control
-
批准号:1351684
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2014
-
负责人:Alexander Olshevsky
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: