CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
批准号:
1652950
负责人:
David Wingate
金额:
$50.97万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-03-01 至 2024-02-29
中文摘要
在强化学习(RL)中,自主智能体(如灾难恢复机器人、自动驾驶汽车或无人驾驶飞行器)在行动时必须同时学习未知环境。深度强化学习是RL的一种变体,它利用深度神经网络(DNN)的能力来学习从传感器提取信息以及如何将这些信息转换为最佳动作。RL和深度学习的结合产生了令人印象深刻的进步,但无论是在数据方面还是在计算方面,它都是昂贵的:典型的代理只有在与环境进行数千万或数亿次交互后才能学习,这使得大多数算法无法在任何地方使用,除非是快速的模拟器。这与人类尝试同样的任务形成了鲜明对比,人类只需几分钟的练习就能表现良好(甚至只需观看另一人几分钟的练习)。从本质上讲,这种类型的机器学习研究构建了供其他人使用的工具。为了将这些工具与理论计算机科学以外的学科联系起来,并改善推广和教育,这项工作整合了一个新的学生交换计划,旨在输出结果和从其他领域引入技术挑战。该计划包括本科生研究机会、竞赛和以项目为导向的课程,重点关注真实世界的系统和数据--所有这些都旨在通过改进技术来激发兴奋并传达对更美好世界的希望。这项工作试图通过解决两个基本问题来改善深度RL:第一,如何减少深度RL所需的数据量和计算量;第二,如何通过整合基于模型的先验知识来提高深度RL解决复杂任务的能力。这一技术策略建立在认知科学的思想基础上,模仿了当前深度RL算法所缺乏的三种关键的人类认知能力:(1)人类天生的建立世界模型的能力,这使得他们(2)通过抽象从先前的经验中转移知识,(3)对自己的不确定性进行明确的推理。为了实现这一点,这项工作结合了两个框架的优点:深度神经网络和贝叶斯模型。DNN提供低级信号处理、灵活且可学习的模型组件以及用于判别性推理的强大构件,而贝叶斯模型在连贯的概率框架中提供关于对象、因果关系和心理理论的高级推理,该框架可以明确地处理不确定性。这些能力是由改进的概率编程框架提供的,该框架支持必要的模型和算法:概率编程允许构建复杂的概率模型,并通过自动推理编译器提供与DNN集成的自然机会。
英文摘要
In reinforcement learning (RL), autonomous agents (such as disaster recovery robots, self-driving cars or unmanned aerial vehicles) must simultaneously learn about an unknown environment while acting in that environment. Deep reinforcement learning is a variant of RL that leverages the power of deep neural networks (DNNs) to learn both to extract information from sensors and how to transform that information into optimal actions. The combination of RL and deep learning has generated impressive advances, but it is expensive, both in terms of data and in terms of computation: typical agents only learn after tens or hundreds of millions of interactions with an environment, making most algorithms unusable in anything but a fast simulator. This stands in stark contrast to humans attempting the same tasks, who can perform well after only a few minutes of practice (or even merely watching another human practice for a few minutes). By its nature, this type of machine learning research builds tools for others to use. To connect these tools with disciplines outside of theoretical computer science and to improve outreach and education, this work integrates a new student exchange program, intended to both export results and to import technical challenges from other fields. The plan is rounded out with a mix of undergraduate research opportunities, competitions, and project-oriented classes focused on real-world systems and data -- all intended to spark excitement and communicate a hope for a better world through improved technology.This work seeks to improve deep RL by addressing two fundamental issues: first, how to reduce the amount of data and computation needed by deep RL, and second, how to improve deep RL's ability to solve complex tasks by incorporating model-based prior knowledge. The technical strategy builds on ideas from cognitive science, mimicing three key human cognitive capabilities lacking in current deep RL algorithms: (1) humans' native ability to build models of the world, which allows them to (2) transfer knowledge from previous experience via abstraction, and (3) reason explicitly about their own uncertainty. To accomplish this, this work combines the strengths of two frameworks: deep neural networks and Bayesian models. The DNNs provide low-level signal processing, flexible and learnable model components, and powerful building blocks for discriminative inference, while the Bayesian models provide high-level reasoning about objects, causality, and theory of mind in a coherent probabilistic framework that can deal explicitly with uncertainty. These capabilities are delivered by improved probabilistic programming frameworks that enable both the necessary models and algorithms: probabilistic programming permits the construction complex probabilistic models, and provides natural opportunities for integration with DNNs through automated inference compilers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
-
批准号:2214708
-
项目类别:Standard Grant
-
资助金额:$101.48万
-
财政年份:2022
-
负责人:David Wingate
-
依托单位:
EAGER: Harnessing Accurate Bias in Large-Scale Language Models
-
批准号:2141680
-
项目类别:Standard Grant
-
资助金额:$27.89万
-
财政年份:2021
-
负责人:David Wingate
-
依托单位:
海外基金