Unifying Recent Advances in Deep Learning with Decision-theoretic Planning for Learned MDPs and POMDPs
Unifying Recent Advances in Deep Learning with Decision-theoretic Planning for Learned MDPs and POMDPs
批准号:
RGPIN-2022-04377
负责人:
Sanner, Scott
金额:
$4.01万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
在从城市交通管理到交互式会话推荐系统的许多复杂的顺序决策问题中,人类很难(如果不是不可能的话)指定这些领域的完整和准确的模型。然而,现代传感的无处不在,从交通摄像头到语音家庭设备,使我们能够从这些复杂的系统中收集大量数据,以学习预测准确的模型,以达到最佳决策理论规划的目的。为此,研究学习模型中的规划有两条线:数据驱动规划(DDP)和基于离线模型的强化学习(MBRL)。虽然DDP和离线MBRL已经取得了重要进展,但它们还没有充分利用规划或深度学习方面的最新进展,这些进展考虑了认知不确定性(即对模型已知内容的信心)、部分可观测性和价值意识(即与决策最低相关的内容),这些对许多现实世界的应用都是至关重要的。我们将在以下研究主题中解决这些不足:主题1--MDP中认知不确定性的深度贝叶斯模型规划:改进规划和模型学习w.r.t.考虑到认知的不确定性,我们将(A)调查贝叶斯深度学习在MDP模型获取方面的最新进展的有效性,(B)为这些贝叶斯深度模型开发稳健的端到端规划方法,以及(C)学习启发式方法以进一步提高规划效率。主题2--深度学习部分观测MDP(POMDP)中的规划:迄今为止,部分可观测性在DDP和离线MBRL中几乎没有受到直接关注。为了解决这一不足,我们将(A)调查用于学习准确POMDP模型的转换器,(B)研究使用复杂观测(例如图像或文本)进行深度贝叶斯信念更新的新方法,以及(C)利用(A)和(B)新的端到端深度学习的POMDP规划技术。主题3--学习的MDP和POMDP中的价值意识:当从丰富但不完整的观测数据中学习时,对于计算效率和样本效率来说,了解什么与预测奖励最不相关是至关重要的。为了解决主题1和主题2的价值意识问题,我们将研究(A)具有认知不确定性的价值意识深度MDP模型和(B)价值意识深度POMDP模型中的规划。这项研究将以两个正在进行的关键应用研究项目为基础:(I)与多伦多大学智能交通系统中心合作的城市交通信号控制的MDP和(Ii)交互式对话推荐系统的POMDP。虽然拟议的研究将从根本上促进深度学习的最新进展与学习的MDP和POMDP中针对各种潜在领域的决策理论规划的统一,但这些特定的应用将作为验证这一研究计划的实践试验台。
英文摘要
In many complex sequential decision-making problems ranging from urban traffic management to interactive conversational recommender systems, it is difficult (if not impossible) for humans to specify complete and accurate models of these domains. However, the ubiquity of modern sensing ranging from traffic cameras to speech-enabled home devices allows us to collect large quantities of data from these complex systems to learn predictively accurate models for purposes of optimal decision-theoretic planning. To this end, there have been two lines of research investigating planning in learned models: data-driven planning (DDP) and offline model-based reinforcement learning (MBRL). While DDP and offline MBRL have made important progress, they have not fully exploited recent advances in planning or deep learning that consider epistemic uncertainty (i.e., confidence in what a model knows), partial observability, and value-awareness (i.e., what is minimally relevant for decision-making) that are critical for many real-world applications. We will address these deficiencies in the following research themes: Theme 1 -- Planning with Deep Bayesian Models of Epistemic Uncertainty in MDPs: To improve planning and model-learning w.r.t. epistemic uncertainty, we will (a) investigate the efficacy of recent advances in Bayesian deep learning for MDP model acquisition, (b) develop robust end-to-end planning methods for these Bayesian deep models, and (c) learn heuristics to further improve planning efficiency. Theme 2 -- Planning in Deep-learned Partially Observed MDPs (POMDPs): To date, partial observability has received little direct attention in DDP and offline MBRL. To address this deficiency, we will (a) investigate transformers for learning accurate POMDP models, (b) investigate novel methods for deep Bayesian belief updating with complex observations (e.g., images or text), and (c) leverage (a) and (b) in novel end-to-end deep-learned POMDP planning techniques. Theme 3 -- Value-awareness in Learned MDPs and POMDPs: When learning from rich, but incomplete observational data, it is critical for both computational and sample efficiency to learn what is minimally relevant for predicting reward. To address value-awareness for Themes 1 and 2, we will investigate planning in (a) value-aware deep MDP models with epistemic uncertainty and (b) value-aware deep POMDP models. The research will be grounded in two key ongoing applied research projects: (i) MDPs for urban traffic signal control in collaboration with the University of Toronto Intelligent Transportation Systems Centre and (ii) POMDPs for interactive conversational recommender systems. While the proposed research will fundamentally contribute to the unification of recent advances in deep learning with decision-theoretic planning in learned MDPs and POMDPs for a variety of potential domains, these specific applications will serve as practical testbeds to validate this research program.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
-
批准号:RGPIN-2016-05705
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.35万
-
财政年份:2021
-
负责人:Sanner, Scott
-
依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
-
批准号:RGPIN-2016-05705
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.35万
-
财政年份:2020
-
负责人:Sanner, Scott
-
依托单位:
Machine learning for residential building HVAC analytics platform
-
批准号:508857-2017
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$1.55万
-
财政年份:2020
-
负责人:Sanner, Scott
-
依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
-
批准号:RGPIN-2016-05705
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.35万
-
财政年份:2019
-
负责人:Sanner, Scott
-
依托单位:
Machine learning for residential building HVAC analytics platform
-
批准号:508857-2017
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$1.55万
-
财政年份:2019
-
负责人:Sanner, Scott
-
依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
-
批准号:RGPIN-2016-05705
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.35万
-
财政年份:2018
-
负责人:Sanner, Scott
-
依托单位:
Machine Learning, Sentiment, and Social Media Analysis for Financial Analytics
-
批准号:531275-2018
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2018
-
负责人:Sanner, Scott
-
依托单位:
Machine learning for residential building HVAC analytics platform
-
批准号:508857-2017
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$1.55万
-
财政年份:2018
-
负责人:Sanner, Scott
-
依托单位:
Machine learning for residential building HVAC analytics platform
-
批准号:508857-2017
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$1.55万
-
财政年份:2017
-
负责人:Sanner, Scott
-
依托单位:
Deep Unsupervised Learning for Network Anomaly Detection
-
批准号:514078-2017
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2017
-
负责人:Sanner, Scott
-
依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
-
批准号:RGPIN-2016-05705
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.35万
-
财政年份:2017
-
负责人:Sanner, Scott
-
依托单位:
Continuous Decision Diagrams for Machine Learning and Decision-theoretic AI Planning
-
批准号:RGPIN-2016-05705
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.35万
-
财政年份:2016
-
负责人:Sanner, Scott
-
依托单位:
Machine learning enabled adaptive user interfaces for enhanced network security
-
批准号:499391-2016
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2016
-
负责人:Sanner, Scott
-
依托单位:
海外基金