课题基金 / 基金详情

Collaborative Research: CIF: Medium: Statistical and Algorithmic Foundations of Distributionally Robust Policy Learning

Collaborative Research: CIF: Medium: Statistical and Algorithmic Foundations of Distributionally Robust Policy Learning
合作研究:CIF:媒介:分布式稳健政策学习的统计和算法基础
批准号:
2312204
负责人:
Jose Blanchet
金额:
$80.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30

项目摘要

项目成果

Jose Blanchet的其他基金

相似基金

相关文献

中文摘要
翻译
高效的数据驱动的政策学习和部署技术正在改变我们社会的许多方面,因为它们在工程、科学和社会应用中具有广泛的适用性。例如,由于能够获得高性能计算,模拟器和数字双胞胎的使用已成为测试和学习复杂优化策略的实用替代方案。因此,在过去的十年里,学术界对这一领域进行了大量的研究。然而,尽管取得了里程碑式的进展,但这一领域的现有工作往往作出一个关键的(隐含的)假设,即培训政策的环境将与部署政策的环境相同。在此假设下学习的策略可能是脆弱的,因为由于模拟器模型规范或环境变化,此假设在实际环境中通常不成立。这个项目的目标是研究统计和算法基础,以便在可能错误指定的生成模型下,在未知环境中开发可证明有效的稳健政策学习。该项目研究了在上下文强盗和强化学习(RL)环境中用于分布式稳健策略学习的全面的统计和算法基础,并在广泛的非参数分布漂移中开发了统计最优和计算高效的算法。这为捕捉与模型无关的环境变化提供了一个强大的框架,但同时也带来了智力挑战,因为未知的最坏情况环境位于无限维空间中。目前的计划开辟了几个基础研究方向,需要新的和原则性的发展。首先,该项目开发了信息理论工具,以了解分布式稳健策略学习的基本学习限制,并表征分布式不确定性如何导致学习困难。此外,该项目还开发了计算高效和统计上最优的估计方案,用于对给定策略进行分布式稳健的性能分析。最后,该项目转化了由于学习分布式稳健政策而在评估中获得的效率收益。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Efficient data-driven policy learning and deployment techniques are transforming many facets of our society as a result of their broad applicability in engineering, scientific and societal applications. Given the access to high-performance computing, the use of simulators and digital twins, for example, have emerged as practical alternatives to test and learn complex optimization policies. As a result, significant scholarly efforts have been devoted to this research area in the past decade. However, despite having made landmark progress, existing work in this area often makes a key (implicit) assumption; namely, that the environment in which the policy is trained will be the same as the environment in which the policy is deployed. Policies learned under this assumption can be fragile, as this assumption often does not hold in practical environments, either due to the simulator model specification or environment shifts. The goal of this project is to study statistical and algorithmic foundations for developing provably efficient robust policy learning in unknown environments, under a possibly misspecified generative model. The project studies comprehensive statistical and algorithmic foundations for distributionally robust policy learning in contextual bandits and reinforcement learning (RL) environments and develops statistically optimal and computationally efficient algorithms across a wide range of non-parametric distributional shifts. These provide a powerful framework for capturing model-agnostic environment changes, but at the same time, pose intellectual challenges as the unknown worst-case environment lies in an infinite-dimensional space. The presented program opens up several fundamental research directions that call for novel and principled developments. First, the project develops information-theoretic tools to understand the fundamental learning limits for distributionally robust policy learning and to characterize how the distributional uncertainty contributes to the difficulty of learning. Additionally, the project develops computationally efficient and statistically optimal estimation schemes for distributionally robust performance analysis of a given policy. Lastly, the project translates the efficiency gains in estimation due to learning a distributionally robust policy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: AMPS: Rare Events in Power Systems: Novel Mathematics, Statistics and Algorithms.
  • 批准号:
    2229011
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2023
  • 负责人:
    Jose Blanchet
  • 依托单位:
DMS-EPSRC: Fast Martingales, Large Deviations, and Randomized Gradients for Heavy-tailed Distributions
  • 批准号:
    2118199
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2021
  • 负责人:
    Jose Blanchet
  • 依托单位:
Robust Wasserstein Profile Inference
  • 批准号:
    1915967
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2019
  • 负责人:
    Jose Blanchet
  • 依托单位:
An Approach to Robust Performance Analysis Using Optimal Transport
  • 批准号:
    1820942
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $24.0万
  • 财政年份:
    2018
  • 负责人:
    Jose Blanchet
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)