An insect-based method for learning landmark reliability using expectation reinforcement in dynamic environments

An insect-based method for learning landmark reliability using expectation reinforcement in dynamic environments
复制标题

DOI:
10.1109/robot.2010.5509935
复制
发表时间:
2010-05
期刊:
2010 IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
Zenon Mathews;P. Verschure;S. Badia
Zenon Mathews;P. Verschure;S. Badia
中科院分区:
其他
文献类型:
--
作者:
Zenon Mathews;P. Verschure;S. Badia

文献摘要

被引文献

相似文献

在未知的动态环境中导航仍然是机器人技术的主要挑战。然而,像沙漠蚁这样的昆虫,它们的计算和记忆能力非常有限,可以非常高效地解决这个问题。因此,了解昆虫导航的潜在神经机制可以告诉我们如何构建更简单但更强大的自主机器人。基于昆虫神经行为学和认知心理学的最新进展,提出了一种动态环境下的地标导航方法。我们的方法使导航器能够使用期望强化方法来学习地标的可靠性。为此,我们实现了一个基于分布式自适应控制框架的实时神经元模型。结果表明,我们的模型能够通过强化其期望来学习地标的稳定性。此外,所提出的机制允许导航器在其期望被违反时以最佳方式恢复其信心。我们还用真实的蚂蚁进行了导航实验,以与我们的模型结果进行比较。所提出的自主导航器的行为与真实蚂蚁的导航行为非常相似。此外,我们的模型将动态环境中的导航解释为一个记忆巩固过程,利用期望及其违反。
Navigation in unknown dynamic environments still remains a major challenge in robotics. Whereas insects like the desert ant with very limited computing and memory capacities solve this task with great efficiency. Thus, the understanding of the underlying neural mechanisms of insect navigation can inform us on how to build simpler yet robust autonomous robots. Based on recent developments in insect neuroethology and cognitive psychology, we propose a method for landmark navigation in dynamic environments. Our method enables the navigator to learn the reliability of landmarks using an expectation reinforcement method. For that end, we implemented a real-time neuronal model based on the Distributed Adaptive Control framework. The results demonstrate that our model is capable of learning the stability of landmarks by reinforcing its expectations. Also, the proposed mechanism allows the navigator to optimally restore its confidence when its expectations are violated. We also perform navigational experiments with real ants to compare with the results of our model. The behavior of the proposed autonomous navigator closely resembles real ant navigational behavior. Moreover, our model explains navigation in dynamic environments as a memory consolidation process, harnessing expectations and their violations.