CAREER: Stochasticity and Resilience in Reinforcement Learning: From Single to Multiple Agents
CAREER: Stochasticity and Resilience in Reinforcement Learning: From Single to Multiple Agents
批准号:
2339794
负责人:
Qiaomin Xie
金额:
$53.29万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-03-01 至 2029-02-28
中文摘要
强化学习(RL)已经成为一种很有前途的数据驱动范式,用于学习控制未知和复杂的系统。它在游戏等模拟环境中取得了令人印象深刻的成功。然而,对于现实世界的工程系统中的应用,现有的RL算法和理论没有解决三个基本挑战:高随机性,长期制度和模型不确定性的脆弱性。这些挑战在具有多个战略代理的系统中加剧。这个CAREER项目的目标是通过解决这些挑战来推进RL的算法和理论基础,并在工程系统中实现高效和弹性的基于RL的控制。该项目将特别侧重于计算机和通信网络的应用,这将指导问题的形成、方法的发展和评价。该项目通过一项教育计划得到加强,该计划旨在为从K-12到大学的学生提供一条途径,以获得RL和广泛的机器学习以及其在工程系统中的应用的经验和培训。该项目还将支持STEM中代表性不足的群体的学生的辅导计划。该项目的研究工作将通过三个技术重点来解决上述挑战。Thrust 1通过统一的变分不等式框架,利用现代马尔可夫链理论的工具,研究RL中出现的各种迭代算法的有限时间收敛性。在第二部分中,我们将开发技术来驯服长期问题中的高随机性,并进一步开发RL算法,以证明学习稳定和接近最优的策略。Thrust 3通过平均场博弈和图子博弈的框架,以及模型不确定性下的鲁棒马尔可夫博弈的博弈理论基础,研究了可扩展的多智能体RL。开发的RL算法将在计算机和通信网络中广泛的决策问题中实施和评估。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Reinforcement Learning (RL) has emerged as a promising data-driven paradigm for learning to control unknown and complex systems. It has achieved impressive success in simulated environments such as games. However, for applications in real-world engineering systems, existing RL algorithms and theory fall short of addressing three fundamental challenges: high stochasticity, long-horizon regimes and vulnerability to model uncertainty. These challenges are exacerbated in systems with multiple strategic agents. The goal of this CAREER project is to advance the algorithmic and theoretical foundations of RL by addressing these challenges, and enable efficient and resilient RL-based control in engineering systems. This project will particularly focus on applications in computer and communication networks, which will guide the problem formulation, methodology development and evaluation. The project is enhanced by an education plan that aims to offer students from K–12 to college a pathway to obtain experience and training in RL and broadly machine learning, as well as in their applications in engineering systems. This project will also support a mentoring program for students fromunderrepresented groups in STEM.The research work in this project will address the aforementioned challenges via three technical thrusts. Thrust 1 studies finite-time convergence of various iterative algorithms that arise in RL through the unified variational inequality framework, by leveraging tools from modern Markov chain theory. In Thrust 2, we will develop techniques to tame the high stochasticity in long-horizon problems, and further develop RL algorithms that provably learn a stable and near-optimal policy. Thrust 3 studies scalable multi-agent RL through the framework of mean-field game and graphon game, as well as the game theoretical foundation of robust Markov games under model uncertainty. The developed RL algorithms will be implemented and evaluated in a broad profile of decision-making problems in computer and communication networks.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Travel: Student Travel Grant for the 2024 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems
-
批准号:2412676
-
项目类别:Standard Grant
-
资助金额:$3.5万
-
财政年份:2024
-
负责人:Qiaomin Xie
-
依托单位:
海外基金