CAREER: Theoretical Foundations of Offline Reinforcement Learning
CAREER: Theoretical Foundations of Offline Reinforcement Learning
批准号:
2141781
负责人:
Nan Jiang
金额:
$50.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-05-01 至 2027-04-30
中文摘要
该奖项的全部或部分资金来自《2021年美国救援计划法案》(公法117-2)。强化学习(RL)是人工智能(AI)的一个子领域,用于解决复杂的决策任务。它在模拟器定义的问题上取得了令人印象深刻的成功,在模拟器定义的问题中,RL代理在虚拟的“在线”环境中通过反复试验学习。然而,很难将这些在线算法应用于现实世界的问题,因为在现实生活中,反复试验往往代价高昂或不可能。例如,个性化医学中的RL代理测试可能会伤害患者的新治疗策略,仅仅是为了收集新的信息,这是不道德的。解决这个问题的一个有前途的范例是离线RL,在那里代理只从历史数据中学习。虽然缺乏与真实环境的直接互动防止了不受欢迎的现实世界后果,但这也给学习带来了重大的技术挑战。该项目旨在开发新的方法来应对这些挑战,并为离线RL提供深入的理论理解,并在使离线RL在机器人、自适应医疗和在线推荐系统等现实应用中取得重大进展。研究进展还将纳入该项目的教育计划,其中包括为代表不足的学生提供建议,开发新课程和一本关于强化学习的专著。该项目的技术目标包括两个方面。第一个重点是模型选择的问题:在训练完成后,我们应该如何在坚持数据集的候选策略之间进行选择?模型选择实现了超参数调整,这是实际机器学习的支柱,但由于问题的多阶段性质,在离线RL中进行超参数调整是出了名的困难。该提案描述了一种很有希望的方法,它建立在研究者最近关于值函数选择的理论工作的基础上。该项目将在理论见解的基础上设计出经验上有效的方法,并解决实际问题,例如不适合的候选函数和复盖面不足的数据。第二个推力思考了线下RL培训的理论基础:在什么条件下才能保证培训的成功?该提案展示了线下RL培训的理论图景,并确定了重要的开放问题和发现新的理论和算法见解的机会。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).Reinforcement learning (RL) is a subarea of Artificial Intelligence (AI) that solves complex decision-making tasks. It has achieved impressive successes in simulator-defined problems, where the RL agent learns via trial-and-error inside a virtual "online" environment. However, it is difficult to apply these online algorithms to real-world problems, as trial-and-error is often expensive or impossible in real life. For example, it is unethical for an RL agent in personalized medicine to test a new treatment strategy that may harm patients, just for the purpose of gathering new information. A promising paradigm to addressing this issue is offline RL, where the agent learns solely from historical data. While the lack of direct interactions with the real environment prevents undesirable real-world consequences, it also gives rise to significant technical challenges in learning. This project aims to develop novel methods to address these challenges and provide a deep theoretical understanding for offline RL, and make significant progress in enabling offline RL in real-life applications such as robotics, adaptive medical treatment, and online recommendation systems. The research development will also be integrated into the project's educational plan, which includes advising underrepresented students and developing new courses and a monograph on reinforcement learning. The technical aims of the project consist of two thrusts. The first thrust focuses on the problem of model selection: after training is completed, how should we select between candidate policies on a holdout dataset? Model selection enables hyperparameter tuning, which is the backbone of practical machine learning, yet it is notoriously difficult in offline RL due to the multi-stage nature of the problem. The proposal describes a promising approach that builds on the investigator's recent theoretical work on value-function selection. The project will devise empirically effective methods based on the theoretical insights and address practical issues such as poorly fitted candidate functions and data with insufficient coverage. The second thrust considers the theoretical foundation of offline RL training: under what conditions can we guarantee the success of training? The proposal lays out the theoretical landscape of offline-RL training, and identifies important open questions and opportunities for discovering novel theoretical and algorithmic insights.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(13)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2307.13332
发表时间:
2023-07
期刊:
影响因子:
--
作者:
[P. Amortila;Nan Jiang;Csaba Szepesvari]
通讯作者:
P. Amortila;Nan Jiang;Csaba Szepesvari
DOI:
10.48550/arxiv.2302.02252
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
作者:
[Audrey Huang;Jinglin Chen;Nan Jiang]
通讯作者:
Audrey Huang;Jinglin Chen;Nan Jiang
DOI:
10.48550/arxiv.2302.11048
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
作者:
[M. Bhardwaj;Tengyang Xie;Byron Boots;Nan Jiang;Ching-An Cheng]
通讯作者:
M. Bhardwaj;Tengyang Xie;Byron Boots;Nan Jiang;Ching-An Cheng
DOI:
10.48550/arxiv.2203.13935
发表时间:
2022-03
期刊:
影响因子:
--
作者:
[Jinglin Chen;Nan Jiang]
通讯作者:
Jinglin Chen;Nan Jiang
DOI:
--
发表时间:
2021-11
期刊:
Applied Bionics and Biomechanics
影响因子:
2.2
作者:
[C. Shi;Masatoshi Uehara;Jiawei Huang;Nan Jiang]
通讯作者:
C. Shi;Masatoshi Uehara;Jiawei Huang;Nan Jiang
共 13 条
CAREER: New Algorithms and Models for Turbulence in Incompressible Fluids
-
批准号:2143331
-
项目类别:Continuing Grant
-
资助金额:$46.28万
-
财政年份:2022
-
负责人:Nan Jiang
-
依托单位:
Probing Local Structural and Chemical Properties of Atomically Thin Two-Dimensional Materials by Optical Scanning Tunneling Microscopy
-
批准号:2211474
-
项目类别:Continuing Grant
-
资助金额:$53.75万
-
财政年份:2022
-
负责人:Nan Jiang
-
依托单位:
Efficient Ensemble Methods for Predictive Fluid Flow Simulations Subject to Uncertainty
-
批准号:2120413
-
项目类别:Standard Grant
-
资助金额:$14.99万
-
财政年份:2021
-
负责人:Nan Jiang
-
依托单位:
CAREER: Probing Chemistry of Surface-Supported Nanostructures at the Angstrom-Scale
-
批准号:1944796
-
项目类别:Continuing Grant
-
资助金额:$68.61万
-
财政年份:2020
-
负责人:Nan Jiang
-
依托单位:
Collaborative Research: Integrated Experimental and Computational Studies for Understanding the Interplay of Photoreactive Materials and Persistent Contaminants
-
批准号:1807465
-
项目类别:Standard Grant
-
资助金额:$23.95万
-
财政年份:2018
-
负责人:Nan Jiang
-
依托单位:
Efficient Ensemble Methods for Predictive Fluid Flow Simulations Subject to Uncertainty
-
批准号:1720001
-
项目类别:Standard Grant
-
资助金额:$14.99万
-
财政年份:2017
-
负责人:Nan Jiang
-
依托单位:
Time-Resolved EELS of Photonic Crystals and Glasses
-
批准号:0603993
-
项目类别:Continuing Grant
-
资助金额:$52.07万
-
财政年份:2006
-
负责人:Nan Jiang
-
依托单位:
海外基金