课题基金 / 基金详情

Unifying Recent Advances in Deep Learning with Decision-theoretic Planning for Learned MDPs and POMDPs

Unifying Recent Advances in Deep Learning with Decision-theoretic Planning for Learned MDPs and POMDPs
将深度学习的最新进展与学习 MDP 和 POMDP 的决策理论规划相结合
批准号:
RGPIN-2022-04377
负责人:
Sanner, Scott
金额:
$4.01万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Sanner, Scott的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
In many complex sequential decision-making problems ranging from urban traffic management to interactive conversational recommender systems, it is difficult (if not impossible) for humans to specify complete and accurate models of these domains. However, the ubiquity of modern sensing ranging from traffic cameras to speech-enabled home devices allows us to collect large quantities of data from these complex systems to learn predictively accurate models for purposes of optimal decision-theoretic planning. To this end, there have been two lines of research investigating planning in learned models: data-driven planning (DDP) and offline model-based reinforcement learning (MBRL). While DDP and offline MBRL have made important progress, they have not fully exploited recent advances in planning or deep learning that consider epistemic uncertainty (i.e., confidence in what a model knows), partial observability, and value-awareness (i.e., what is minimally relevant for decision-making) that are critical for many real-world applications. We will address these deficiencies in the following research themes: Theme 1 -- Planning with Deep Bayesian Models of Epistemic Uncertainty in MDPs: To improve planning and model-learning w.r.t. epistemic uncertainty, we will (a) investigate the efficacy of recent advances in Bayesian deep learning for MDP model acquisition, (b) develop robust end-to-end planning methods for these Bayesian deep models, and (c) learn heuristics to further improve planning efficiency. Theme 2 -- Planning in Deep-learned Partially Observed MDPs (POMDPs): To date, partial observability has received little direct attention in DDP and offline MBRL. To address this deficiency, we will (a) investigate transformers for learning accurate POMDP models, (b) investigate novel methods for deep Bayesian belief updating with complex observations (e.g., images or text), and (c) leverage (a) and (b) in novel end-to-end deep-learned POMDP planning techniques. Theme 3 -- Value-awareness in Learned MDPs and POMDPs: When learning from rich, but incomplete observational data, it is critical for both computational and sample efficiency to learn what is minimally relevant for predicting reward. To address value-awareness for Themes 1 and 2, we will investigate planning in (a) value-aware deep MDP models with epistemic uncertainty and (b) value-aware deep POMDP models. The research will be grounded in two key ongoing applied research projects: (i) MDPs for urban traffic signal control in collaboration with the University of Toronto Intelligent Transportation Systems Centre and (ii) POMDPs for interactive conversational recommender systems. While the proposed research will fundamentally contribute to the unification of recent advances in deep learning with decision-theoretic planning in learned MDPs and POMDPs for a variety of potential domains, these specific applications will serve as practical testbeds to validate this research program.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
  • 批准号:
    RGPIN-2016-05705
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2021
  • 负责人:
    Sanner, Scott
  • 依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
  • 批准号:
    RGPIN-2016-05705
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2020
  • 负责人:
    Sanner, Scott
  • 依托单位:
Machine learning for residential building HVAC analytics platform
  • 批准号:
    508857-2017
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $1.55万
  • 财政年份:
    2020
  • 负责人:
    Sanner, Scott
  • 依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
  • 批准号:
    RGPIN-2016-05705
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2019
  • 负责人:
    Sanner, Scott
  • 依托单位:
海外基金