课题基金 / 基金详情

CAREER: Robust Reinforcement Learning Under Model Uncertainty: Algorithms and Fundamental Limits

CAREER: Robust Reinforcement Learning Under Model Uncertainty: Algorithms and Fundamental Limits
职业:模型不确定性下的鲁棒强化学习:算法和基本限制
批准号:
2337375
负责人:
Shaofeng Zou
金额:
$52.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-09-01 至 2029-08-31

项目摘要

项目成果

Shaofeng Zou的其他基金

相似基金

相关文献

中文摘要
翻译
现有的强化学习(RL)方法通常假设学习的策略将部署在与其训练时相同的环境中。例如,由于对抗性扰动、模拟器和真实世界应用之间的建模误差、非平稳环境和有限的训练数据量,这种假设在实践中经常被违反。训练和测试环境之间的差异导致模型不匹配,这导致性能显著下降,并限制了RL在关键领域的适用性,例如医疗保健、关键基础设施、交通系统和智能城市。为了应对上述挑战,已经做出了值得注意的努力,以开发分发上强大的RL方法。这个职业项目旨在提高分布式健壮RL的基本算法和理论极限。该项目的研究成果有望推动稳健RL的算法和理论边界,并将提供可证明的收敛、高效和极小极大最优的稳健RL算法。该项目将对特殊教育、智能交通系统、无线通信网络、电力系统和无人机网络等各个领域的顺序决策的理论和实践产生重大影响。本项目中的活动将提供具体的原则和设计指南,以实现面对模型不确定性的稳健性。将研究工作整合到教育和外展活动中,将以K-12教育工作者、研究生、本科生和代表性不足的学生为目标,努力(I)为K-12教育工作者举办人工智能(AI)夏令营;(Ii)布法罗日研讨会;(Iii)课程开发;(Iv)学生监督。研究工作围绕三个免费的推动力组织:(I)推力A专注于在长期平均回报标准下为分布稳健的RL发展理论和算法基础。(Ii)推力B侧重于从没有主动数据获取和探索的离线数据集中建立一个用于学习(稳健)策略的分布稳健性的统一框架,并进一步揭示其基本限制;(Iii)推力C侧重于在约束下的稳健风险学习的构造性方法和基本限制,即在模型不确定性下优化回报的同时保证约束。这个项目将发展对健壮RL、极小极大最优健壮RL算法的基本理解,以及新的技术收敛和复杂性分析。研究成果将显著提高RL算法的稳健性,并将引起广泛的社区的兴趣,例如机器学习、统计学、信息论、网络、通信、电力和教育。这项拟议的工作还将促进这些研究界新的跨学科研究方向。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Existing reinforcement learning (RL) approaches usually assume that a learned policy will be deployed in the same environment as the one it was trained in. Such an assumption is often violated in practice, due to e.g., adversarial perturbations, modeling error between simulator and real-world applications, non-stationary environment, and limited amount of training data. The discrepancy between the training and test environments gives rise to a model mismatch, which lead to a notable decline in performance and restrict the suitability of RL in crucial domains, e.g., healthcare, critical infrastructure, transportation systems, and smart cities. To address the above challenge, there have been noteworthy efforts to develop distributionally robust RL approaches. This CAREER project aims to advance the fundamental algorithmic and theoretic limits of distributionally robust RL. The research outcome of this project holds the promise to push the algorithmic and theoretical boundaries of robust RL, and will deliver provably convergent, efficient and minimax optimal robust RL algorithms. The project will have a significant impact on theory and practice of sequential decision making in various domains, e.g., special education, intelligent transportation system, wireless communication networks, power systems and drone networks. The activities in this project will provide concrete principles and design guidelines to achieve robustness in face of model uncertainty. The integration of research work into education and outreach will target K-12 educators, graduate, undergraduate and underrepresented students with efforts on (i) Artificial Intelligence (AI) summer camp for K-12 educators; (ii) Buffalo Day workshop; (iii) curriculum development; (iv) student supervision.The research efforts are organized around three complimentary thrusts: (i) Thrust A focuses on developing theoretical and algorithmic foundations for distributionally robust RL under the long-term average-reward criterion. (ii) Thrust B focuses on developing a unified framework of distributional robustness for learning (robust) policies from offline dataset without active data acquisition and exploration, and further uncovering their fundamental limits; (iii) Thrust C focuses on constructive approaches and fundamental limits of robust RL under constraints, i.e., optimizing reward while simultaneously guaranteeing constraints under model uncertainty. This project will develop fundamental understandings of robust RL, minimax optimal robust RL algorithms and novel technical convergence and complexity analyses. The research outcome will significantly improve the robustness of RL algorithms and will be of interest to a broad range of communities, e.g., machine learning, statistics, information theory, networking, communication, power, and education. The proposed work will also foster new interdisciplinary research directions across these research communities.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CIF: Medium: Emerging Directions in Robust Learning and Inference
  • 批准号:
    2106560
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $37.47万
  • 财政年份:
    2021
  • 负责人:
    Shaofeng Zou
  • 依托单位:
CCSS: Collaborative Research: Quickest Threat Detection in Adversarial Sensor Networks
  • 批准号:
    2112693
  • 项目类别:
    Standard Grant
  • 资助金额:
    $21.7万
  • 财政年份:
    2021
  • 负责人:
    Shaofeng Zou
  • 依托单位:
CRII: CIF: Dynamic Network Event Detection with Time-Series Data
  • 批准号:
    1948165
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.49万
  • 财政年份:
    2020
  • 负责人:
    Shaofeng Zou
  • 依托单位:
CIF: Small: Reinforcement Learning with Function Approximation: Convergent Algorithms and Finite-sample Analysis
  • 批准号:
    2007783
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.0万
  • 财政年份:
    2020
  • 负责人:
    Shaofeng Zou
  • 依托单位:
国内基金
海外基金
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
  • 批准号:
    70601028
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    7.0万元
  • 批准年份:
    2006
  • 负责人:
    王明征
  • 依托单位:
心理紧张和应力影响下Robust语音识别方法研究
  • 批准号:
    60085001
  • 项目类别:
    专项基金项目
  • 资助金额:
    14.0万元
  • 批准年份:
    2000
  • 负责人:
    韩纪庆
  • 依托单位:
ROBUST语音识别方法的研究
  • 批准号:
    69075008
  • 项目类别:
    面上项目
  • 资助金额:
    3.5万元
  • 批准年份:
    1990
  • 负责人:
    高雨青
  • 依托单位:
改进型ROBUST序贯检测技术
  • 批准号:
    68671030
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1986
  • 负责人:
    刘有恒
  • 依托单位: