CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
批准号:
1652950
负责人:
David Wingate
金额:
$50.97万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-03-01 至 2024-02-29
中文摘要
在强化学习(RL)中,自主代理(如灾难恢复机器人、自动驾驶汽车或无人驾驶飞行器)必须同时学习未知环境,同时在该环境中行动。深度强化学习是强化学习的一种变体,它利用深度神经网络(dnn)的力量来学习从传感器中提取信息以及如何将这些信息转化为最佳行为。强化学习和深度学习的结合产生了令人印象深刻的进步,但它在数据和计算方面都是昂贵的:典型的代理只有在与环境进行数千万或数亿次交互后才能学习,这使得大多数算法只能在快速模拟器中使用。这与尝试同样任务的人类形成了鲜明的对比,人类只需练习几分钟(甚至只是看另一个人练习几分钟)就能表现得很好。就其本质而言,这种类型的机器学习研究构建了供他人使用的工具。为了将这些工具与理论计算机科学以外的学科联系起来,并改善推广和教育,这项工作整合了一个新的学生交换计划,旨在输出结果并从其他领域引入技术挑战。该计划将本科生研究机会、竞赛和以项目为导向的课程结合起来,重点关注现实世界的系统和数据——所有这些都旨在激发人们的兴奋,并通过改进的技术传达对更美好世界的希望。这项工作旨在通过解决两个基本问题来改进深度强化学习:第一,如何减少深度强化学习所需的数据量和计算量;第二,如何通过结合基于模型的先验知识来提高深度强化学习解决复杂任务的能力。该技术策略建立在认知科学思想的基础上,模仿了当前深度强化学习算法中缺乏的三个关键的人类认知能力:(1)人类建立世界模型的天生能力,这使得他们能够(2)通过抽象从以前的经验中转移知识,以及(3)明确地对自己的不确定性进行推理。为了做到这一点,这项工作结合了两个框架的优势:深度神经网络和贝叶斯模型。dnn提供低级信号处理、灵活且可学习的模型组件,以及用于判别推理的强大构建块,而贝叶斯模型在连贯的概率框架中提供关于对象、因果关系和心智理论的高级推理,可以明确地处理不确定性。这些功能由改进的概率编程框架提供,该框架支持必要的模型和算法:概率编程允许构建复杂的概率模型,并通过自动推理编译器提供与dnn集成的自然机会。
英文摘要
In reinforcement learning (RL), autonomous agents (such as disaster recovery robots, self-driving cars or unmanned aerial vehicles) must simultaneously learn about an unknown environment while acting in that environment. Deep reinforcement learning is a variant of RL that leverages the power of deep neural networks (DNNs) to learn both to extract information from sensors and how to transform that information into optimal actions. The combination of RL and deep learning has generated impressive advances, but it is expensive, both in terms of data and in terms of computation: typical agents only learn after tens or hundreds of millions of interactions with an environment, making most algorithms unusable in anything but a fast simulator. This stands in stark contrast to humans attempting the same tasks, who can perform well after only a few minutes of practice (or even merely watching another human practice for a few minutes). By its nature, this type of machine learning research builds tools for others to use. To connect these tools with disciplines outside of theoretical computer science and to improve outreach and education, this work integrates a new student exchange program, intended to both export results and to import technical challenges from other fields. The plan is rounded out with a mix of undergraduate research opportunities, competitions, and project-oriented classes focused on real-world systems and data -- all intended to spark excitement and communicate a hope for a better world through improved technology.This work seeks to improve deep RL by addressing two fundamental issues: first, how to reduce the amount of data and computation needed by deep RL, and second, how to improve deep RL's ability to solve complex tasks by incorporating model-based prior knowledge. The technical strategy builds on ideas from cognitive science, mimicing three key human cognitive capabilities lacking in current deep RL algorithms: (1) humans' native ability to build models of the world, which allows them to (2) transfer knowledge from previous experience via abstraction, and (3) reason explicitly about their own uncertainty. To accomplish this, this work combines the strengths of two frameworks: deep neural networks and Bayesian models. The DNNs provide low-level signal processing, flexible and learnable model components, and powerful building blocks for discriminative inference, while the Bayesian models provide high-level reasoning about objects, causality, and theory of mind in a coherent probabilistic framework that can deal explicitly with uncertainty. These capabilities are delivered by improved probabilistic programming frameworks that enable both the necessary models and algorithms: probabilistic programming permits the construction complex probabilistic models, and provides natural opportunities for integration with DNNs through automated inference compilers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
-
批准号:2214708
-
项目类别:Standard Grant
-
资助金额:$101.48万
-
财政年份:2022
-
负责人:David Wingate
-
依托单位:
EAGER: Harnessing Accurate Bias in Large-Scale Language Models
-
批准号:2141680
-
项目类别:Standard Grant
-
资助金额:$27.89万
-
财政年份:2021
-
负责人:David Wingate
-
依托单位:
海外基金