CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
批准号:
1652950
负责人:
David Wingate
金额:
$50.97万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-03-01 至 2024-02-29
中文摘要
在强化学习(RL)中,自主智能体(如灾难恢复机器人、自动驾驶汽车或无人机)必须在未知环境中行动时同时学习该环境。 深度强化学习是RL的一种变体,它利用深度神经网络(DNN)的能力来学习从传感器中提取信息以及如何将这些信息转换为最佳动作。 RL和深度学习的结合产生了令人印象深刻的进步,但无论是在数据方面还是在计算方面,它都是昂贵的:典型的代理只有在与环境进行数千万或数亿次交互后才能学习,这使得大多数算法除了快速模拟器之外无法使用。 这与尝试相同任务的人类形成了鲜明的对比,人类只需几分钟的练习(甚至只是看另一个人练习几分钟)就可以表现得很好。就其本质而言,这种类型的机器学习研究为其他人提供了工具。为了将这些工具与理论计算机科学之外的学科联系起来,并改善推广和教育,这项工作整合了一个新的学生交流计划,旨在输出结果并从其他领域导入技术挑战。该计划通过本科生研究机会、竞赛和以现实世界系统和数据为重点的项目导向课程的组合来完善-所有这些都旨在激发兴奋并通过改进的技术传达对更美好世界的希望。这项工作旨在通过解决两个基本问题来改善深度RL:首先,如何减少深度RL所需的数据量和计算量,其次,如何通过结合基于模型的先验知识来提高深度RL解决复杂任务的能力。 该技术策略建立在认知科学的思想基础上,模仿当前深度RL算法缺乏的三种关键人类认知能力:(1)人类构建世界模型的能力,这使他们能够(2)通过抽象从以前的经验中转移知识,以及(3)明确地推理自己的不确定性。为了实现这一目标,这项工作结合了两个框架的优势:深度神经网络和贝叶斯模型。 DNN提供了低级信号处理,灵活且可学习的模型组件以及用于判别式推理的强大构建块,而贝叶斯模型在一个连贯的概率框架中提供了关于对象,因果关系和心理理论的高级推理,可以明确地处理不确定性。 这些功能是通过改进的概率编程框架提供的,这些框架支持必要的模型和算法:概率编程允许构建复杂的概率模型,并通过自动推理编译器提供与DNN集成的自然机会。
英文摘要
In reinforcement learning (RL), autonomous agents (such as disaster recovery robots, self-driving cars or unmanned aerial vehicles) must simultaneously learn about an unknown environment while acting in that environment. Deep reinforcement learning is a variant of RL that leverages the power of deep neural networks (DNNs) to learn both to extract information from sensors and how to transform that information into optimal actions. The combination of RL and deep learning has generated impressive advances, but it is expensive, both in terms of data and in terms of computation: typical agents only learn after tens or hundreds of millions of interactions with an environment, making most algorithms unusable in anything but a fast simulator. This stands in stark contrast to humans attempting the same tasks, who can perform well after only a few minutes of practice (or even merely watching another human practice for a few minutes). By its nature, this type of machine learning research builds tools for others to use. To connect these tools with disciplines outside of theoretical computer science and to improve outreach and education, this work integrates a new student exchange program, intended to both export results and to import technical challenges from other fields. The plan is rounded out with a mix of undergraduate research opportunities, competitions, and project-oriented classes focused on real-world systems and data -- all intended to spark excitement and communicate a hope for a better world through improved technology.This work seeks to improve deep RL by addressing two fundamental issues: first, how to reduce the amount of data and computation needed by deep RL, and second, how to improve deep RL's ability to solve complex tasks by incorporating model-based prior knowledge. The technical strategy builds on ideas from cognitive science, mimicing three key human cognitive capabilities lacking in current deep RL algorithms: (1) humans' native ability to build models of the world, which allows them to (2) transfer knowledge from previous experience via abstraction, and (3) reason explicitly about their own uncertainty. To accomplish this, this work combines the strengths of two frameworks: deep neural networks and Bayesian models. The DNNs provide low-level signal processing, flexible and learnable model components, and powerful building blocks for discriminative inference, while the Bayesian models provide high-level reasoning about objects, causality, and theory of mind in a coherent probabilistic framework that can deal explicitly with uncertainty. These capabilities are delivered by improved probabilistic programming frameworks that enable both the necessary models and algorithms: probabilistic programming permits the construction complex probabilistic models, and provides natural opportunities for integration with DNNs through automated inference compilers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
-
批准号:2214708
-
项目类别:Standard Grant
-
资助金额:$101.48万
-
财政年份:2022
-
负责人:David Wingate
-
依托单位:
EAGER: Harnessing Accurate Bias in Large-Scale Language Models
-
批准号:2141680
-
项目类别:Standard Grant
-
资助金额:$27.89万
-
财政年份:2021
-
负责人:David Wingate
-
依托单位:
海外基金