课题基金 / 基金详情

RI: SMALL: Robust Reinforcement Learning Using Bayesian Models

RI: SMALL: Robust Reinforcement Learning Using Bayesian Models
RI:小:使用贝叶斯模型的鲁棒强化学习
批准号:
1815275
负责人:
Marek Petrik
金额:
$43.78万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-15 至 2023-07-31

项目摘要

项目成果

Marek Petrik的其他基金

相似基金

相关文献

中文摘要
翻译
基于数据的决策比依赖启发式或经验法则更可取。然而,有效地使用数据可能具有挑战性。在农业或医学等领域,数据集通常很小,有偏差,而且有噪声。例如,减少农药使用的全部效果取决于天气,而对产量的影响可能要到收获时才能知道。减少农药的使用可以降低成本,并提供生态和消费者利益,但使用太少很容易导致作物歉收和重大经济损失。数据可用性有限和故障成本高的双重问题在制造、维护甚至机器人技术中也很常见。因为大多数现有的强化学习方法都假设了大型数据集,利益相关者通常会忽略数据驱动的方法,而依赖启发式来做出看似安全但并非最优的决策。这项研究为数据驱动的决策开发了新的强大方法,即使在数据有限的情况下,也可以推荐安全的良好行动。新的强化学习方法使用先验领域知识来估计可能结果的置信度,以防止在预测不正确时发生灾难性故障。这些方法的实际可行性是在使用历史数据建议改进果园农药时间表的问题上进行测试的,并传播给从业人员。本研究针对1)有限或昂贵的数据和2)高失败成本的强化学习问题。当错误的决策造成巨大损失、伤害或死亡时,对政策质量的信心比其最优性差距更重要。在强化学习中计算高置信度策略是困难的。即使是很小的误差也会通过正反馈循环和协变量移位迅速累积。因此,需要更强大的方法来说服从业者从数据中受益,而不是依赖于启发式。该项目将鲁棒优化与基于模型的强化学习相结合,以计算出能够抵抗数据错误的良好策略。鲁棒优化在许多领域取得了成功,但很难与强化学习一起使用。它需要一个合理的不确定性水平模型,即所谓的模糊集,来适当地平衡解决方案。质量和信心。即使对于鲁棒优化专家来说,在序列决策问题中手工构造好的模糊集也是非常困难的。本研究探讨了一种新的数据驱动的稳健强化学习贝叶斯方法。它将分层贝叶斯模型与鲁棒优化相结合,利用强大的分层建模技术,同时避免了贝叶斯强化学习通常带来的计算复杂性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Basing decisions on data is preferable to relying on heuristics or rules of thumb. Using data effectively, however, can be challenging. In domains like agriculture or medicine, datasets are usually small, biased, and noisy. For instance, the full effects of reduced pesticide applications depend on the weather and the impacts on yield may not be known until the harvest. Reducing pesticide applications reduces costs and provides ecological and consumer benefits, but using too little of it can easily cause a crop failure and significant financial losses. These dual problems of limited data availability and a high cost of failure are also common in manufacturing, maintenance, and even robotics. Because most existing reinforcement learning methods assume large datasets, stakeholders often dismiss data-driven methods and rely on heuristics to make decisions that are apparently safe but quite sub-optimal. This research develops new robust methods for data-driven decision making that can recommend good actions that are also safe even when data is limited. The new reinforcement learning methods use prior domain knowledge to estimate the confidence in possible outcomes to prevent catastrophic failure when predictions are incorrect. The practical viability of these methods is tested on the problem of using historical data to recommending improved pesticide schedules for fruit orchards and is disseminated to practitioners.This research targets reinforcement learning problems with 1) limited or expensive data and 2) a high cost of failure. When bad decisions cause large losses, injury, or death, then having confidence in a policy's quality is more important than its optimality gap. Computing high-confidence policies in reinforcement learning is difficult. Even small errors can quickly accumulate through positive feedback loops and covariate shift. Therefore, more robust methods are needed to convince practitioners to benefit from data instead of relying on heuristics. The project combines robust optimization with model-based reinforcement learning to compute good policies that are resistant to data errors. Robust optimization has achieved successes in many areas but can be difficult to use with reinforcement learning. It requires a model of plausible uncertainty levels, so-called ambiguity sets, to properly balance solution?s quality and confidence. Constructing good ambiguity sets manually in sequential decision problems is very difficult even for robust optimization experts. This research investigates a new data-driven Bayesian approach to robust reinforcement learning. It combines hierarchical Bayesian models with robust optimization to leverage powerful hierarchical modeling techniques while avoiding the computational complexity often associated with Bayesian reinforcement learning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(15)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2019-02
期刊: ArXiv
影响因子: --
作者: [Marek Petrik;R. Russel]
通讯作者: Marek Petrik;R. Russel
Optimizing Percentile Criterion Using Robust MDPs
使用稳健的 MDP 优化百分位数标准
DOI: --
发表时间: 2021
期刊: Proceedings of Machine Learning Research
影响因子: --
作者: [Bahram Behzadian, Reazul Hasan]
通讯作者: Bahram Behzadian, Reazul Hasan
Fast Algorithms for L-infinity constrained S-rectangular Robust MDPs
L-无穷大约束 S-矩形鲁棒 MDP 的快速算法
DOI: --
发表时间: 2021
期刊: Neural Information Processing Systems
影响因子: --
作者: [Bahram Behzadian, Marek Petrik]
通讯作者: Bahram Behzadian, Marek Petrik
DOI: --
发表时间: 2020-07
期刊: ArXiv
影响因子: --
作者: [Daniel S. Brown;S. Niekum;Marek Petrik]
通讯作者: Daniel S. Brown;S. Niekum;Marek Petrik
共 13 条
    CAREER: Soft-robust Methods for Offline Reinforcement Learning
    • 批准号:
      2144601
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $57.59万
    • 财政年份:
      2022
    • 负责人:
      Marek Petrik
    • 依托单位:
    III: Small: Robust Reinforcement Learning for Invasive Species Management
    • 批准号:
      1717368
    • 项目类别:
      Standard Grant
    • 资助金额:
      $49.73万
    • 财政年份:
      2017
    • 负责人:
      Marek Petrik
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: