RI: Small: Towards Provably Efficient Representation Learning in Reinforcement Learning via Rich Function Approximation
RI: Small: Towards Provably Efficient Representation Learning in Reinforcement Learning via Rich Function Approximation
批准号:
2154711
负责人:
Wen Sun
金额:
$38.46万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-10-01 至 2025-09-30
中文摘要
强化学习使人工智能系统能够自我学习。虽然今天的强化学习系统在某些任务(如国际象棋)上的经验表现优于人类,但这些系统通常依赖于极端数量的数据和计算资源。这使得它们不适合数据昂贵的现实应用。此外,这些系统中使用的强化学习算法通常没有任何性能保证,例如算法需要多少数据点才能以高置信度解决任务,这也限制了它们在安全关键应用中的使用。该项目的主要新奇将是开发新的强化学习算法,这些算法可以有效地学习,使用尽可能少的训练数据点。开发高效的强化学习算法可以将这些系统的应用扩展到收集数据成本高昂的现实应用中。例如,在自动驾驶系统中,开发的技术将有可能使自动驾驶汽车通过减少错误来更快地适应新的道路条件。在针对视障人士的个性化导航系统中,使用高效强化学习算法训练的系统可以在学习过程的早期阶段通过高质量的交互与用户互动,该项目旨在弥合强化学习理论与实践之间的差距,通过开发计算和统计上有效的算法,规模马尔可夫决策过程,其中数据是高维和复杂的。该项目提出的关键创新是通过将表征学习纳入强化学习框架来打开黑盒。表示学习方法允许算法从高维和非结构化数据中提取紧凑的信息,并仅使用紧凑的表示进行推理和决策,从而大大提高了样本和计算效率。两个主要目标是:(1)如何学习用于反事实强化学习的表示,其中学习者只能访问静态数据集,并且没有能力与环境进一步交互;(2)如何将表示学习,探索和开发集成到在线强化学习设置中,其中代理需要主动与环境交互以获取数据。除了算法和强化学习表示学习理论的发展外,本项目还提出设计能够适应最终用户的个性化语音导航系统,其中,样本高效的离线和在线强化学习在快速安全的适应中发挥着重要作用。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准。
英文摘要
Reinforcement Learning enables artificial intelligence systems to learn by themselves. While today’s reinforcement Learning systems can empirically outperform humans on some tasks (such as chess), these systems often rely on an extreme amount of data and computation resources. This makes them not suitable for real-world applications where data are expensive. Also, reinforcement learning algorithms used in these systems often do not have any performance guarantees, such as how many data points the algorithm needs in order to solve the task with high confidence, which also limits their usage in safety critical applications. The main novelty of this project will be the development of new reinforcement learning algorithms that can learn efficiently, using as few training data points as possible. The development of efficient reinforcement learning algorithms can expand the applications of these systems to real-world applications where data are expensive to collect. For example, in autonomous driving systems, the developed technologies would have the potential to enable self-driving cars to adapt to new road conditions faster by making fewer mistakes. In personalized navigation systems for visually impaired people, systems trained with efficient reinforcement learning algorithms can engage with users via high-quality interactions at an early stage of the learning process, thus positively influence the user experience.The project aims to bridge the gap between reinforcement learning theory and practice by developing computationally and statistically efficient algorithms for large-scale Markov Decision Processes where data are high-dimensional and complex. The key innovation proposed in this project is to open the black box by incorporating representation learning into the reinforcement learning framework. The representation learning approach allows algorithms to extract compact information from high dimensional and unstructured data, and perform reasoning and decision making only using the compact representation — thus vastly improving the sample and computation efficiency. Two main thrusts are: (1) how to learn representations for counterfactual reinforcement learning where the learner only has access to a static dataset and has no ability to further interact with the environment; (2) how to integrate representation learning, exploration, and exploitation in the online reinforcement learning setting where the agent needs to actively interact with the environment for data acquisition. In addition to the algorithms and reinforcement learning representation learning theory development, this project proposes to design personalized voice navigation systems that can adapt to end-users, where sample efficient offline and online reinforcement learning plays an important role in fast and safe adaptation.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.48550/arxiv.2302.04441
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
作者:
[Yihan Du;Longbo Huang;Wen Sun]
通讯作者:
Yihan Du;Longbo Huang;Wen Sun
DOI:
10.48550/arxiv.2205.14571
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
作者:
[Alekh Agarwal;Yuda Song;Wen Sun;Kaiwen Wang;Mengdi Wang;Xuezhou Zhang]
通讯作者:
Alekh Agarwal;Yuda Song;Wen Sun;Kaiwen Wang;Mengdi Wang;Xuezhou Zhang
DOI:
10.48550/arxiv.2210.06718
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
作者:
[Yuda Song;Yi Zhou;Ayush Sekhari;J. Bagnell;A. Krishnamurthy;Wen Sun]
通讯作者:
Yuda Song;Yi Zhou;Ayush Sekhari;J. Bagnell;A. Krishnamurthy;Wen Sun
CAREER: Towards Real-world Reinforcement Learning
-
批准号:2339395
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2024
-
负责人:Wen Sun
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: