CIF: Small: Accelerating Stochastic Approximation for Optimization and Reinforcement Learning
CIF: Small: Accelerating Stochastic Approximation for Optimization and Reinforcement Learning
批准号:
2306023
负责人:
Sean Meyn
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2026-06-30
中文摘要
本计画系关于递归演算法之设计与分析,此演算法在工程与电脑科学上有广泛之应用。递归算法在ChatGPT等机器学习系统中发挥着至关重要的作用,这些系统依赖于大量数据进行训练。强化学习是一个有许多著名例子的领域,它利用递归算法来训练计算机程序;其中最著名的例子包括在围棋和国际象棋等游戏中表现出色的计算机程序。训练被解释为“学习”最佳响应(例如,下一步棋)基于观察(棋盘的当前布局)。虽然随机逼近被认为是递归算法的数学模型,并在学习的数学理论中发挥了重要作用,但支持理论并没有跟上经验的成功。在强化学习中,通常不确定训练是否会成功或需要多少训练。 沿着为算法学习创造新基础的基础研究,该研究项目还涉及研究生指导,通过在线视频讲座传播新的和现有的研究成果,并通过研究人员组织的认知和控制研讨会进行传播,该研讨会每年在佛罗里达大学举行,吸引来自美国各地和国外的演讲者。 技术将被开发,以确保稳定性和加速收敛的随机近似算法的瞬态和方差。算法设计的新方法将包括基于常微分方程方法的技术,马尔可夫过程的最新理论,以及基于准随机探索的学习方法。 算法设计中的大部分工作都归结为反馈控制问题,最初是在连续时间内提出的,以利用非线性控制和稳定性理论的概念。 一个显著的例子是牛顿-拉夫逊流,它在温和的假设下是全局收敛的。 一个可靠的“算法反馈法”在连续时间,然后转化为一个可靠的和有效的算法在离散时间实现。一般理论将在两个特定应用领域中发展:强化学习和无梯度优化。强化学习提出了最大的挑战,因为到目前为止,除了非常特殊的情况外,几乎没有理论可以建立这些递归算法的稳定性。 此外,在最近的工作中,研究人员和他的学生已经表明,马尔可夫记忆可以导致非常缓慢的收敛,即使当算法被优化时;在这种情况下,有必要改变算法目标,而不会对算法提供的最终解决方案的质量产生负面影响。 在强化学习的情况下,主要目标是有效地学习用于决策的有效规则(即,政策)。幸运的是,在选择适合学习最佳策略的标准方面有很大的自由。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的知识价值和更广泛的影响审查标准进行评估来支持。
英文摘要
This project concerns the design and analysis of recursive algorithms, which have broad applications in engineering and computer science. Recursive algorithms play a crucial role in machine learning systems like ChatGPT, which rely on large amounts of data for training. Reinforcement learning, a field with numerous famous examples, utilizes recursive algorithms for training computer programs; among the most famous examples include computer programs that excel in games such as GO and chess. Training is interpreted as "learning" optimal responses (e.g., the next move) based on observations (the current configuration of a chessboard). While stochastic approximation is recognized as a mathematical model for recursive algorithms and plays a major role in the mathematical theory of learning, the supporting theory has not kept pace with empirical success. In reinforcement learning, it is often uncertain if training will be successful or how much training is required. Along with fundamental research to create new foundations for algorithmic learning, the research project also involves graduate student mentoring, dissemination of new and existing research results through online video lectures, and also dissemination through the Workshop on Cognition and Control organized by the investigator, which is held annually at the University of Florida attracting speakers from across the U.S. and abroad. Techniques will be developed to ensure stability and accelerate convergence of stochastic approximation algorithms in terms of transients and variance. New approaches to algorithm design will include techniques based on ordinary differential equation methods, recent theory of Markov processes, and approaches to learning based on quasi-random exploration. Much of the work in algorithm design reduces to a feedback control problem, initially posed in continuous time to leverage concepts from nonlinear control and stability theory. A remarkable example is the Newton-Raphson flow which is globally convergent under mild assumptions. A dependable "algorithmic feedback law" in continuous time is then translated into a reliable and efficient algorithm implemented in discrete time. The general theory will be developed within two specific application areas: reinforcement learning and gradient-free optimization. Reinforcement learning presents the greatest challenge because, to-date, there is little theory available to establish the stability of these recursive algorithms outside of very special cases. Moreover, in recent work the investigator with his students have shown that Markovian memory can result in very slow convergence, even when the algorithm is optimized; in such cases it is necessary to change the algorithmic goal without negatively impacting the quality of the final solution delivered by the algorithm. In the case of reinforcement learning the primary objective is to efficiently learn an effective rule for decision making (i.e., a policy). Fortunately, there is great freedom in choosing a criterion of fit for learning the best policy within a given class.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Characterizing capacity of controllable DERs to provide energy storage service to the power grid
-
批准号:2122313
-
项目类别:Standard Grant
-
资助金额:$35.97万
-
财政年份:2021
-
负责人:Sean Meyn
-
依托单位:
Reinforcement Learning and Kullback-Leibler Stochastic Optimal Control for Complex Networks
-
批准号:1935389
-
项目类别:Standard Grant
-
资助金额:$38.0万
-
财政年份:2019
-
负责人:Sean Meyn
-
依托单位:
Distributed Control for Demand Dispatch: The Creation of Virtual Energy Storage from Flexible Loads
-
批准号:1609131
-
项目类别:Standard Grant
-
资助金额:$38.0万
-
财政年份:2016
-
负责人:Sean Meyn
-
依托单位:
CPS:Medium:Collaborative Research: Smart Power Systems of the Future: Foundations for Understanding Volatility and Improving Operational Reliability
-
批准号:1259040
-
项目类别:Standard Grant
-
资助金额:$69.72万
-
财政年份:2012
-
负责人:Sean Meyn
-
依托单位:
CPS:Medium:Collaborative Research: Smart Power Systems of the Future: Foundations for Understanding Volatility and Improving Operational Reliability
-
批准号:1135598
-
项目类别:Standard Grant
-
资助金额:$70.0万
-
财政年份:2011
-
负责人:Sean Meyn
-
依托单位:
Robust Inference and Communication: Theory, Algorithms and Performance Analysis
-
批准号:0729031
-
项目类别:Standard Grant
-
资助金额:$38.0万
-
财政年份:2007
-
负责人:Sean Meyn
-
依托单位:
Control Techniques for Complex Networks
-
批准号:0523620
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Sean Meyn
-
依托单位:
Visualization & Optimization Techniques For Analysis and Design of Complex Systems
-
批准号:0217836
-
项目类别:Standard Grant
-
资助金额:$19.5万
-
财政年份:2002
-
负责人:Sean Meyn
-
依托单位:
US-India Workshop: Learning, Adaptation, and Optimization, Kerala, India, December 2000
-
批准号:0079744
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2000
-
负责人:Sean Meyn
-
依托单位:
Optimization and Performance Evaluation of Network Models
-
批准号:9972957
-
项目类别:Standard Grant
-
资助金额:$23.36万
-
财政年份:1999
-
负责人:Sean Meyn
-
依托单位:
Systems Design and Analysis: Stability, Performance and Robustness
-
批准号:9403742
-
项目类别:Standard Grant
-
资助金额:$22.5万
-
财政年份:1994
-
负责人:Sean Meyn
-
依托单位:
Research Initiation Award: Techniques for Design and Analysis of Short Memory Stochastic Adaptive Control Algorithms
-
批准号:8910088
-
项目类别:Standard Grant
-
资助金额:$4.5万
-
财政年份:1989
-
负责人:Sean Meyn
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: