RI: Small: Towards Optimal and Adaptive Reinforcement Learning with Offline Data and Limited Adaptivity
RI: Small: Towards Optimal and Adaptive Reinforcement Learning with Offline Data and Limited Adaptivity
批准号:
2007117
负责人:
Yu-Xiang Wang
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30
中文摘要
强化学习(RL)是机器学习中发展最快的研究领域之一。基于强化学习的技术最近在人工智能领域取得了几项突破,比如在围棋比赛中击败了人类冠军。然而,RL在现实生活问题中的应用仍然有限,即使在已经收集了大量数据的领域也是如此。问题的关键在于,大多数现有的强化学习方法都需要一个代理与之交互的环境,但在现实应用中,很少有可能访问这样的环境——部署一个通过试错来学习的算法可能会有严重的法律、道德和安全问题。这个项目旨在通过开发从离线数据中学习的算法来解决这个难题。这项研究的结果可以显著减少在现实生活中使用强化学习技术解决顺序决策问题的开销,比如电力传输、个性化医疗、科学发现、计算机网络和公共政策。该项目侧重于两种环境,旨在解决上述环境有限的挑战。在第一种设置中,只向代理提供与环境交互记录的历史数据。在第二种设置中,代理只能改变它与环境的交互方式几次。研究人员将开发数学理论来描述问题的难度,并确保所开发的算法在使用尽可能少的资源(数据、能源、计算)的意义上是鲁棒和最佳的。利用边缘重要性采样、均匀收敛和批量探索等技术,该项目将推广“打破地平线诅咒”的最新工作路线,以允许函数近似,并为离线和低自适应强化学习建立急需的统计学习理论。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Reinforcement learning (RL) is one of the fastest-growing research areas in machine learning. RL-based techniques have led to several recent breakthroughs in artificial intelligence, such as beating human champions in the game of Go. The application of RL to real life problems, however, remains limited, even in areas where a large amount of data has already been collected. The crux of the problem is that most existing RL methods require an environment for the agent to interact with, but in real-life applications, it is rarely possible to have access to such an environment — deploying an algorithm that learns by trial-and-errors may have serious legal, ethical and safety issues. This project aims to address this conundrum by developing algorithms that learn from offline data. The outcome of the research could significantly reduce the overhead of using RL techniques in real-life sequential decision-making problems such as those in power transmission, personalized medicine, scientific discoveries, computer networking and public policy.The project focuses on two settings that aim at addressing the aforementioned challenge of limited access to an environment. In the first setting, the agent is given only the historical data from logged interactions with the environment. In the second setting, the agent is able to change how it interacts with the environment only a few times. The investigators will develop mathematical theory that describes the difficulty of the problem and ensures that the developed algorithms are robust and optimal in the sense that they use the least possible resources (data, energy, computation). Using techniques such as marginalized importance sampling, uniform convergence and batched exploration, the project will generalize the recent line of work in ``breaking the curse of horizon'' to allow function approximations and establish the much-needed statistical learning theory for offline and low-adaptive reinforcement learning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(16)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
作者:
[Dan Qiao;Ming Yin;Ming Min;Yu-Xiang Wang]
通讯作者:
Dan Qiao;Ming Yin;Ming Min;Yu-Xiang Wang
DOI:
10.48550/arxiv.2302.13252
发表时间:
2023-02
期刊:
影响因子:
--
作者:
[Chong Liu;Ming Yin;Yu-Xiang Wang]
通讯作者:
Chong Liu;Ming Yin;Yu-Xiang Wang
DOI:
10.48550/arxiv.2210.00701
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
作者:
[Dan Qiao;Yu-Xiang Wang]
通讯作者:
Dan Qiao;Yu-Xiang Wang
DOI:
10.48550/arxiv.2211.15956
发表时间:
2022-11
期刊:
影响因子:
--
作者:
[Jiachen Li;Edwin Zhang;Ming Yin;Qinxun Bai;Yu-Xiang Wang;William Yang Wang]
通讯作者:
Jiachen Li;Edwin Zhang;Ming Yin;Qinxun Bai;Yu-Xiang Wang;William Yang Wang
DOI:
10.48550/arxiv.2212.04680
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
作者:
[Dan Qiao;Yu-Xiang Wang]
通讯作者:
Dan Qiao;Yu-Xiang Wang
共 15 条
CAREER: Exact Optimal and Data-Adaptive Algorithms and Tools for Differential Privacy
-
批准号:2048091
-
项目类别:Continuing Grant
-
资助金额:$49.92万
-
财政年份:2021
-
负责人:Yu-Xiang Wang
-
依托单位:
Collaborative Research: SCALE MoDL: Adaptivity of Deep Neural Networks
-
批准号:2134214
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2021
-
负责人:Yu-Xiang Wang
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: