Reducing Computational Cost During Robot Navigation and Human-Robot Interaction with a Human-Inspired Reinforcement Learning Architecture

Reducing Computational Cost During Robot Navigation and Human-Robot Interaction with a Human-Inspired Reinforcement Learning Architecture
复制标题

DOI:
10.1007/s12369-022-00942-6
复制
发表时间:
2022-11-08
影响因子:
4.7
通讯作者:
Khamassi, Mehdi
Khamassi, Mehdi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dromnelle, Remi;Renaudo, Erwan;Khamassi, Mehdi

文献摘要

被引文献

相似文献

我们提出了一种新的神经启发的强化学习架构,用于机器人在社交和非社交场景中的在线学习和决策。我们的目标是从人类根据自身表现的变化动态和自主地调整行为的方式中获得灵感,同时最大限度地减少认知努力。遵循计算神经科学原理,该架构结合了基于模型(MB)和无模型(MF)的强化学习(RL)。这里的主要新奇在于仲裁与元控制器选择当前的学习策略,根据效率和计算成本之间的权衡。MB策略构建了行动的长期影响模型,并使用该模型通过动态规划进行决策,从而以高计算成本为代价灵活适应任务变化。MF策略灵活性较低,但成本也低1000倍,并通过观察MB决策进行学习。我们测试的架构在三个实验:一个导航任务在真实的环境中的任务变化(墙壁配置的变化,目标位置的变化);一个模拟的对象操作任务下的人类教学信号;和一个模拟的人机合作任务,整理对象在桌子上。我们表明,我们的人类启发的策略协调方法,使机器人保持最佳的性能方面的奖励和计算成本相比,MB专家单独,实现了最佳的性能,但具有最高的计算成本。我们还表明,该方法使得有可能科普环境的突然变化,目标的变化或人类合作伙伴的行为在交互任务的变化。执行这些实验的机器人,无论是真实的还是虚拟的,都使用相同的参数集,从而显示了该方法的通用性。
We present a new neuro-inspired reinforcement learning architecture for robot online learning and decision-making during both social and non-social scenarios. The goal is to take inspiration from the way humans dynamically and autonomously adapt their behavior according to variations in their own performance while minimizing cognitive effort. Following computational neuroscience principles, the architecture combines model-based (MB) and model-free (MF) reinforcement learning (RL). The main novelty here consists in arbitrating with a meta-controller which selects the current learning strategy according to a trade-off between efficiency and computational cost. The MB strategy, which builds a model of the long-term effects of actions and uses this model to decide through dynamic programming, enables flexible adaptation to task changes at the expense of high computation costs. The MF strategy is less flexible but also 1000 times less costly, and learns by observation of MB decisions. We test the architecture in three experiments: a navigation task in a real environment with task changes (wall configuration changes, goal location changes); a simulated object manipulation task under human teaching signals; and a simulated human-robot cooperation task to tidy up objects on a table. We show that our human-inspired strategy coordination method enables the robot to maintain an optimal performance in terms of reward and computational cost compared to an MB expert alone, which achieves the best performance but has the highest computational cost. We also show that the method makes it possible to cope with sudden changes in the environment, goal changes or changes in the behavior of the human partner during interaction tasks. The robots that performed these experiments, whether real or virtual, all used the same set of parameters, thus showing the generality of the method.