Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments

Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments
复制标题

DOI:
10.1109/icra40945.2020.9196922
复制
发表时间:
2020-03
期刊:
2020 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Stewart Jamieson;J. How;Yogesh A. Girdhar
Stewart Jamieson;J. How;Yogesh A. Girdhar
中科院分区:
其他
文献类型:
--
作者:
Stewart Jamieson;J. How;Yogesh A. Girdhar

文献摘要

相似文献

我们提出了一种新的POMDP问题制定的机器人,必须自主决定去哪里收集新的和科学相关的图像与人类操作员的通信能力有限。从这个公式中,我们得到的约束条件和设计原则的观察模型,奖励模型,和这样的机器人的通信策略,探索技术来处理非常高维的观察空间和相关的训练数据的稀缺性。我们引入了一种新的主动奖励学习策略的基础上,使查询,以帮助机器人最大限度地减少路径“遗憾”在线,并通过模拟评估其适用于自主视觉探索。我们证明,在一些带宽有限的环境中,这种新的后悔为基础的标准,使机器人探险家收集高达17%以上的奖励每使命比下一个最好的标准。
We present a novel POMDP problem formulation for a robot that must autonomously decide where to go to collect new and scientifically relevant images given a limited ability to communicate with its human operator. From this formulation we derive constraints and design principles for the observation model, reward model, and communication strategy of such a robot, exploring techniques to deal with the very high-dimensional observation space and scarcity of relevant training data. We introduce a novel active reward learning strategy based on making queries to help the robot minimize path "regret" online, and evaluate it for suitability in autonomous visual exploration through simulations. We demonstrate that, in some bandwidth-limited environments, this novel regret-based criterion enables the robotic explorer to collect up to 17% more reward per mission than the next-best criterion.