Learning Task-Driven Control Policies via Information Bottlenecks

Learning Task-Driven Control Policies via Information Bottlenecks
复制标题

DOI:
10.15607/rss.2020.xvi.101
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Vincent Pacelli;Anirudha Majumdar
Vincent Pacelli;Anirudha Majumdar
中科院分区:
其他
文献类型:
--
作者:
Vincent Pacelli;Anirudha Majumdar

文献摘要

相似文献

本文提出了一种强化学习方法,用于为配备有丰富感觉模态(例如,视觉或深度)。标准的强化学习算法通常会产生将控制动作与整个系统的状态和丰富的传感器观测紧密耦合的策略。因此,产生的策略通常可能对状态或观察结果中与任务无关的部分的变化敏感(例如,改变背景颜色)。相比之下,我们在这里提出的方法学习创建一个任务驱动的表示,用于计算控制动作。从形式上讲,这是通过导出一个策略梯度式算法来实现的,该算法在状态和任务驱动表示之间创建了一个信息瓶颈;这将动作约束为仅依赖于任务相关信息。我们在多个示例的一组完整的模拟结果中展示了我们的方法,包括利用深度图像的抓取任务和利用RGB图像的接球任务。与标准的政策梯度方法的比较表明,我们的算法产生的任务驱动的政策往往是更强大的传感器噪声和任务无关的环境变化。
This paper presents a reinforcement learning approach to synthesizing task-driven control policies for robotic systems equipped with rich sensory modalities (e.g., vision or depth). Standard reinforcement learning algorithms typically produce policies that tightly couple control actions to the entirety of the system's state and rich sensor observations. As a consequence, the resulting policies can often be sensitive to changes in task-irrelevant portions of the state or observations (e.g., changing background colors). In contrast, the approach we present here learns to create a task-driven representation that is used to compute control actions. Formally, this is achieved by deriving a policy gradient-style algorithm that creates an information bottleneck between the states and the task-driven representation; this constrains actions to only depend on task-relevant information. We demonstrate our approach in a thorough set of simulation results on multiple examples including a grasping task that utilizes depth images and a ball-catching task that utilizes RGB images. Comparisons with a standard policy gradient approach demonstrate that the task-driven policies produced by our algorithm are often significantly more robust to sensor noise and task-irrelevant changes in the environment.