Belief-Grounded Networks for Accelerated Robot Learning under Partial Observability

Belief-Grounded Networks for Accelerated Robot Learning under Partial Observability
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Hai V. Nguyen;Brett Daley;Xinchao Song;Chris Amato;Robert W. Platt
Hai V. Nguyen;Brett Daley;Xinchao Song;Chris Amato;Robert W. Platt
中科院分区:
其他
文献类型:
--
作者:
Hai V. Nguyen;Brett Daley;Xinchao Song;Chris Amato;Robert W. Platt

文献摘要

相似文献

许多重要的机器人问题是部分可观察的,因为单个视觉或力反馈测量不足以重建状态。标准的方法包括学习一个关于信念或观察-行动历史的政策。然而,这两种方法都有缺点;在线跟踪信念的成本很高,而且很难直接通过历史来学习策略。我们提出了一种在部分可观测性下进行策略学习的方法,称为信念接地网络(BGN),其中辅助信念重建损失激励神经网络简洁地总结其输入历史。由于所产生的策略是历史而不是信念的函数,因此它可以在运行时轻松执行。我们比较BGN与几个基线的经典基准任务,以及三个新的机器人触摸感应任务。BGN优于所有其他测试方法,其学习策略在转移到物理机器人上时工作良好。
Many important robotics problems are partially observable in the sense that a single visual or force-feedback measurement is insufficient to reconstruct the state. Standard approaches involve learning a policy over beliefs or observation-action histories. However, both of these have drawbacks; it is expensive to track the belief online, and it is hard to learn policies directly over histories. We propose a method for policy learning under partial observability called the Belief-Grounded Network (BGN) in which an auxiliary belief-reconstruction loss incentivizes a neural network to concisely summarize its input history. Since the resulting policy is a function of the history rather than the belief, it can be executed easily at runtime. We compare BGN against several baselines on classic benchmark tasks as well as three novel robotic touch-sensing tasks. BGN outperforms all other tested methods and its learned policies work well when transferred onto a physical robot.