CAREER: Foundations of Reinforcement Learning under Partial Observability
CAREER: Foundations of Reinforcement Learning under Partial Observability
批准号:
2239297
负责人:
Chi Jin
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-01 至 2028-07-31
中文摘要
现代人工智能的许多挑战可以归结为部分可观测性下的强化学习(RL)问题,在该问题中,智能体学习做出一系列决策,尽管缺乏关于决策做出的时刻到时刻的完整信息。这种部分可观测RL的自然应用包括机器人、自主驾驶、不完全信息博弈、部分信息下的资源分配、行星探测、医疗诊断系统等。因此,Porl一直是运筹学、控制和机器学习领域的一个重要课题。虽然最近社区见证了强化学习理论在完全可观测环境中的突破性进展,但我们对学习在部分可观测系统中行为的理解仍然非常有限。部分可观测性给RL在建模、算法设计和理论分析方面带来了一系列新的独特挑战。解决这些挑战将在学术界、工业界和社会上产生深远的影响,可以应用现代RL。本项目旨在识别和应对这些独特的挑战,建立坚实的理论基础,并为POL设计新的可靠和高效的算法。具体地说,这一提议将从三个渐进的角度来研究波尔。推力1考虑部分可观测马尔可夫决策过程(POMDP)模型下的基本表格设置。这一努力的主要目标是确定允许在统计或计算上有效学习的关键结构条件,并解决推断潜在状态和探索的核心挑战。推力2涉及具有大量状态和观察的现代POIL,其中必须采用函数近似来近似模型、价值函数或策略。我们将在更一般的预测状态表示(PSR)模型下研究这些问题,并在存在函数逼近的情况下开发有效的学习结果。推力3在部分可观测马尔可夫博弈(POMG)模型下,研究了多代理环境下的POL。我们将设计有效的算法来学习POMG中的各种均衡,并解决多机构和分散算法设计带来的独特挑战。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
A wide range of modern artificial intelligence challenges can be cast as Reinforcement Learning (RL) problems under partial observability, in which agents learn to make a sequence of decisions despite lacking complete information about the moment-to-moment situation in which decisions are made. Natural applications of this kind of Partially Observable RL (PORL) include robotics, autonomous driving, imperfect information games, resource allocation under partial information, planetary exploration, medical diagnostic systems. As such, PORL has been an important topic in operation research, control, and machine learning. While the community recently witnessed a surge of breakthroughs in reinforcement learning theory in fully observable environments, our understanding of learning to act in partially observable systems remains very limited. Partial observability brings a new series of unique challenges to RL in modeling, algorithm design, and theoretical analyses. Resolving these challenges will have far-reaching impacts in academia, industry and society where modern RL can be applied.This project aims to identify and attack these unique challenges, establish solid theoretical foundations, and design new reliable and efficient algorithms for PORL. Concretely, this proposal will study PORL in three progressive thrusts. Thrust 1 considers the basic tabular setup, under the model of Partially Observable Markov Decision Processes (POMDPs). The main objective in this thrust is to identify the key structural conditions that permit statistically or computationally efficient learning, and to address the core challenges of inferring latent states and exploration. Thrust 2 concerns modern PORL with an enormous number of states and observations, where function approximation must be deployed to approximate the models, the value functions, or the policies. We will investigate these problems under a more general model of Predictive State Representations (PSRs) and develop efficient learning results in the presence of function approximation. Thrust 3 investigates PORL in the multiagent setting, under the model of Partially Observable Markov Games (POMGs). We will design efficient algorithms for learning various equilibria in POMGs and address the unique challenges arising from multiagency and the design of decentralized algorithms.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Frameworks: hpcGPT: Enhancing Computing Center User Support with HPC-enriched Generative AI
-
批准号:2411299
-
项目类别:Standard Grant
-
资助金额:$36.0万
-
财政年份:2024
-
负责人:Chi Jin
-
依托单位:
RI: Medium: Provable Reinforcement Learning with Function Approximation and Neural Networks
-
批准号:2107304
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2021
-
负责人:Chi Jin
-
依托单位:
海外基金