课题基金 / 基金详情

Unifying Recent Advances in Deep Learning with Decision-theoretic Planning for Learned MDPs and POMDPs

Unifying Recent Advances in Deep Learning with Decision-theoretic Planning for Learned MDPs and POMDPs
将深度学习的最新进展与学习 MDP 和 POMDP 的决策理论规划相结合
批准号:
RGPIN-2022-04377
负责人:
Sanner, Scott
金额:
$4.01万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Sanner, Scott的其他基金

相似基金

相关文献

中文摘要
翻译
在从城市交通管理到交互式会话推荐系统的许多复杂的顺序决策问题中,人类很难(如果不是不可能的话)指定这些领域的完整且准确的模型。然而,从交通摄像头到支持语音的家用设备等现代传感技术的普遍存在,使我们能够从这些复杂的系统中收集大量数据,以学习预测准确的模型,从而实现最佳决策理论规划。为此,有两条研究路线研究学习模型中的规划:数据驱动规划(DDP)和基于离线模型的强化学习(MBRL)。虽然 DDP 和离线 MBRL 取得了重要进展,但它们尚未充分利用规划或深度学习方面的最新进展,这些进展考虑了对于许多现实世界应用至关重要的认知不确定性(即对模型所知道的内容的信心)、部分可观察性和价值意识(即与决策最不相关的内容)。我们将在以下研究主题中解决这些缺陷:主题 1——使用 MDP 中认知不确定性的深贝叶斯模型进行规划:改进规划和模型学习。为了解决认知不确定性,我们将(a)研究贝叶斯深度学习的最新进展对于 MDP 模型获取的有效性,(b)为这些贝叶斯深度模型开发强大的端到端规划方法,以及(c)学习启发式方法以进一步提高规划效率。主题 2——深度学习部分可观测 MDP (POMDP) 中的规划:迄今为止,部分可观测性在 DDP 和离线 MBRL 中几乎没有受到直接关注。为了解决这一缺陷,我们将(a)研究用于学习准确的 POMDP 模型的转换器,(b)研究使用复杂观察(例如图像或文本)进行深度贝叶斯信念更新的新方法,以及(c)利用新颖的端到端深度学习 POMDP 规划技术中的(a)和(b)。主题 3——学习 MDP 和 POMDP 中的价值意识:当从丰富但不完整的观察数据中学习时,了解与预测奖励最不相关的内容对于计算和样本效率至关重要。为了解决主题 1 和 2 的价值意识,我们将研究 (a) 具有认知不确定性的价值意识深度 MDP 模型和 (b) 价值意识深度 POMDP 模型中的规划。该研究将基于两个正在进行的关键应用研究项目:(i)与多伦多大学智能交通系统中心合作的城市交通信号控制 MDP 和(ii)交互式会话推荐系统的 POMDP。虽然拟议的研究将从根本上有助于将深度学习的最新进展与各种潜在领域的学习 MDP 和 POMDP 中的决策理论规划相结合,但这些特定应用程序将作为验证该研究计划的实际测试平台。
英文摘要
In many complex sequential decision-making problems ranging from urban traffic management to interactive conversational recommender systems, it is difficult (if not impossible) for humans to specify complete and accurate models of these domains. However, the ubiquity of modern sensing ranging from traffic cameras to speech-enabled home devices allows us to collect large quantities of data from these complex systems to learn predictively accurate models for purposes of optimal decision-theoretic planning. To this end, there have been two lines of research investigating planning in learned models: data-driven planning (DDP) and offline model-based reinforcement learning (MBRL). While DDP and offline MBRL have made important progress, they have not fully exploited recent advances in planning or deep learning that consider epistemic uncertainty (i.e., confidence in what a model knows), partial observability, and value-awareness (i.e., what is minimally relevant for decision-making) that are critical for many real-world applications. We will address these deficiencies in the following research themes: Theme 1 -- Planning with Deep Bayesian Models of Epistemic Uncertainty in MDPs: To improve planning and model-learning w.r.t. epistemic uncertainty, we will (a) investigate the efficacy of recent advances in Bayesian deep learning for MDP model acquisition, (b) develop robust end-to-end planning methods for these Bayesian deep models, and (c) learn heuristics to further improve planning efficiency. Theme 2 -- Planning in Deep-learned Partially Observed MDPs (POMDPs): To date, partial observability has received little direct attention in DDP and offline MBRL. To address this deficiency, we will (a) investigate transformers for learning accurate POMDP models, (b) investigate novel methods for deep Bayesian belief updating with complex observations (e.g., images or text), and (c) leverage (a) and (b) in novel end-to-end deep-learned POMDP planning techniques. Theme 3 -- Value-awareness in Learned MDPs and POMDPs: When learning from rich, but incomplete observational data, it is critical for both computational and sample efficiency to learn what is minimally relevant for predicting reward. To address value-awareness for Themes 1 and 2, we will investigate planning in (a) value-aware deep MDP models with epistemic uncertainty and (b) value-aware deep POMDP models. The research will be grounded in two key ongoing applied research projects: (i) MDPs for urban traffic signal control in collaboration with the University of Toronto Intelligent Transportation Systems Centre and (ii) POMDPs for interactive conversational recommender systems. While the proposed research will fundamentally contribute to the unification of recent advances in deep learning with decision-theoretic planning in learned MDPs and POMDPs for a variety of potential domains, these specific applications will serve as practical testbeds to validate this research program.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
  • 批准号:
    RGPIN-2016-05705
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2021
  • 负责人:
    Sanner, Scott
  • 依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
  • 批准号:
    RGPIN-2016-05705
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2020
  • 负责人:
    Sanner, Scott
  • 依托单位:
Machine learning for residential building HVAC analytics platform
  • 批准号:
    508857-2017
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $1.55万
  • 财政年份:
    2020
  • 负责人:
    Sanner, Scott
  • 依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
  • 批准号:
    RGPIN-2016-05705
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2019
  • 负责人:
    Sanner, Scott
  • 依托单位:
海外基金