课题基金 / 基金详情

RI: Small: Towards Optimal and Adaptive Reinforcement Learning with Offline Data and Limited Adaptivity

RI: Small: Towards Optimal and Adaptive Reinforcement Learning with Offline Data and Limited Adaptivity
RI:小型:利用离线数据和有限的适应性实现最优和自适应强化学习
批准号:
2007117
负责人:
Yu-Xiang Wang
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30

项目摘要

项目成果

Yu-Xiang Wang的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Reinforcement learning (RL) is one of the fastest-growing research areas in machine learning. RL-based techniques have led to several recent breakthroughs in artificial intelligence, such as beating human champions in the game of Go. The application of RL to real life problems, however, remains limited, even in areas where a large amount of data has already been collected. The crux of the problem is that most existing RL methods require an environment for the agent to interact with, but in real-life applications, it is rarely possible to have access to such an environment — deploying an algorithm that learns by trial-and-errors may have serious legal, ethical and safety issues. This project aims to address this conundrum by developing algorithms that learn from offline data. The outcome of the research could significantly reduce the overhead of using RL techniques in real-life sequential decision-making problems such as those in power transmission, personalized medicine, scientific discoveries, computer networking and public policy.The project focuses on two settings that aim at addressing the aforementioned challenge of limited access to an environment. In the first setting, the agent is given only the historical data from logged interactions with the environment. In the second setting, the agent is able to change how it interacts with the environment only a few times. The investigators will develop mathematical theory that describes the difficulty of the problem and ensures that the developed algorithms are robust and optimal in the sense that they use the least possible resources (data, energy, computation). Using techniques such as marginalized importance sampling, uniform convergence and batched exploration, the project will generalize the recent line of work in ``breaking the curse of horizon'' to allow function approximations and establish the much-needed statistical learning theory for offline and low-adaptive reinforcement learning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(16)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2022-02
期刊: ArXiv
影响因子: --
作者: [Dan Qiao;Ming Yin;Ming Min;Yu-Xiang Wang]
通讯作者: Dan Qiao;Ming Yin;Ming Min;Yu-Xiang Wang
DOI: 10.48550/arxiv.2302.13252
发表时间: 2023-02
期刊:
影响因子: --
作者: [Chong Liu;Ming Yin;Yu-Xiang Wang]
通讯作者: Chong Liu;Ming Yin;Yu-Xiang Wang
DOI: 10.48550/arxiv.2210.00701
发表时间: 2022-10
期刊: ArXiv
影响因子: --
作者: [Dan Qiao;Yu-Xiang Wang]
通讯作者: Dan Qiao;Yu-Xiang Wang
DOI: 10.48550/arxiv.2211.15956
发表时间: 2022-11
期刊:
影响因子: --
作者: [Jiachen Li;Edwin Zhang;Ming Yin;Qinxun Bai;Yu-Xiang Wang;William Yang Wang]
通讯作者: Jiachen Li;Edwin Zhang;Ming Yin;Qinxun Bai;Yu-Xiang Wang;William Yang Wang
15
    CAREER: Exact Optimal and Data-Adaptive Algorithms and Tools for Differential Privacy
    Collaborative Research: SCALE MoDL: Adaptivity of Deep Neural Networks
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: