课题基金 / 基金详情

Robust Decision-Aware Model-based Reinforcement Learning

Robust Decision-Aware Model-based Reinforcement Learning
基于鲁棒决策感知模型的强化学习
批准号:
RGPIN-2021-03701
负责人:
Farahmand, Amirmassoud
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Farahmand, Amirmassoud的其他基金

相似基金

相关文献

中文摘要
翻译
强化学习(RL)是设计一个与环境交互并自适应地提高其长期性能的代理的问题。许多复杂的现实世界决策问题都可以表示为RL问题。示例应用包括混合动力汽车的能源管理系统,医疗保健的动态治疗制度,以及机器人、金融等领域的许多其他应用。RL是人工智能的核心,具有对我们的经济和社会产生巨大影响的潜力,可以说比任何其他机器学习领域都更重要。尽管取得了这些成功,但RL作为一种技术还没有为大多数现实世界的应用程序做好准备。困难的一个主要来源是RL代理的高样本复杂性。样本复杂性指的是实现某一性能级别所需的交互(或数据点)数量。在执行良好之前需要太多样本的RL代理不适合现实世界的应用,在现实世界中,获得新样本通常是昂贵和耗时的。基于模型的RL(Model-Based RL,MBRL)是一种很有前途的方法,可以用来设计样本高效的代理,用于与现实世界交互的数量不能很大的问题。MBRL的基本思想是学习一个环境模型,然后在内部模拟器中使用该模型来规划一个好的策略,即选择动作的策略。这可能会提高代理的样本复杂性。然而,这取决于学习真实世界的准确模型。传统的基于学习环境预测模型的模型学习方法有一个很大的缺陷。它是基于这样一种信念,即一个准确的预测因素就足以进行规划。一个经常被忽视的事实是,任何模型都不可能完全准确,现实世界和模型之间总是存在一些误差。对于我们的模型来说,现实世界有时太复杂了。在我的研究计划中,我的建议是重新考虑我们应该如何做MBRL。试图学习与潜在决策问题无关的复杂动态是没有意义的。传统的模型学习方法不能区分环境中与决策相关和无关的方面,因此将模型的能力浪费在不必要的细节上。这个研究计划的基本思想是,与其试图学习一个能够很好地预测环境的模型,不如只学习与决策问题相关的方面。这个研究项目的科学影响是,它打开并探索了一种非正统的思维方式,即代理人应该如何了解其环境。我期待我的研究团队在这一方向上的进展为基于模型的RL的未来提供理论和基础基础。我还希望它能带来样本效率高的RL代理,可以用于现实世界的应用。
英文摘要
Reinforcement learning (RL) is the problem of designing an agent that interacts with its environment and adaptively improves its long-term performance. Many complex real-world decision-making problems can be formulated as an RL problem. Example applications include energy management systems for hybrid cars, dynamic treatment regimes in healthcare, and many others in robotics, finance, etc. RL is at the core of AI and has the potential of having a huge impact on our economy and society, arguably more so than any other area of machine learning. Despite these successes, RL as a technology is not ready for most real-world applications. A major source of difficulty is the high sample complexity of RL agents. Sample complexity refers to the number of interactions (or data points) required to achieve a certain level of performance. An RL agent that requires too many samples before performing well is unsuitable for real-world applications, in which obtaining new samples is often costly and time consuming. Model-based RL (MBRL) is a promising approach to design sample-efficient agents for problems where the number of interactions with the real-world cannot be very large. The basic idea of MBRL is to learn a model of the environment, and then use the model in an internal simulator to plan a good policy, i.e., the strategy to select actions. This may improve the sample complexity of the agent. This is contingent, however, on learning an accurate model of the real-world. The conventional approach to model learning, which is based on learning a good predictive model of the environment, has an important shortcoming. It is based on the belief that an accurate predictor is sufficient for planning. The often-unnoticed fact is that no model can be completely accurate, and there are always some errors between the real-world and the model. The real-world is sometimes too complex for our models. What I suggest in my research program is to rethink how we should do MBRL. Trying to learn complex dynamics that are irrelevant to the underlying decision problem is pointless. A conventional model learning approach cannot discriminate between decision-relevant and irrelevant aspects of the environment, and hence wastes the capacity of a model on unnecessary detail. The fundamental idea of this research program is that instead of trying to learn a model that is a good predictor of the environment, one should only learn about the aspects that are relevant to the decision problem. The scientific impact of this research program is that it opens up and explores an unorthodox way of thinking about how an agent should learn about its environment. I expect my research team's progress on this direction to provide the theoretical and foundational groundwork for the future of model-based RL. I also expect that it leads to sample-efficient RL agents that can be used for real-world applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Robust Decision-Aware Model-based Reinforcement Learning
  • 批准号:
    DGECR-2021-00419
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2021
  • 负责人:
    Farahmand, Amirmassoud
  • 依托单位:
Robust Decision-Aware Model-based Reinforcement Learning
  • 批准号:
    RGPIN-2021-03701
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2021
  • 负责人:
    Farahmand, Amirmassoud
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis