Object-Centric Scene Representations Using Active Inference

Object-Centric Scene Representations Using Active Inference
复制标题

DOI:
10.1162/neco_a_01637
复制
发表时间:
2024-03-21
期刊:
影响因子:
2.9
通讯作者:
Dhoedt,Bart
Dhoedt,Bart
中科院分区:
计算机科学4区
文献类型:
--
作者:
Van de Maele,Toon;Verbelen,Tim;Dhoedt,Bart

文献摘要

被引文献

相似文献

从原始传感数据中表示场景及其组成对象是使机器人能够与其环境交互的核心能力。在这封信中,我们提出了一种新的场景理解方法,利用以对象为中心的生成模型,使代理推断对象类别和姿态在allocentric参考框架使用主动推理,神经启发的框架行动和感知。为了评估主动视觉代理的行为,我们还提出了一个新的基准,给定一个特定对象的目标视点,代理需要找到最佳匹配的视点,给定一个工作空间,在3D中随机定位的对象。我们证明,我们的主动推理代理能够平衡认知觅食和目标驱动行为,并且在成功率方面定量优于监督学习和强化学习基线两倍以上。
Representing a scene and its constituent objects from raw sensory data is a core ability for enabling robots to interact with their environment. In this letter, we propose a novel approach for scene understanding, leveraging an object-centric generative model that enables an agent to infer object category and pose in an allocentric reference frame using active inference, a neuro-inspired framework for action and perception. For evaluating the behavior of an active vision agent, we also propose a new benchmark where, given a target viewpoint of a particular object, the agent needs to find the best matching viewpoint given a workspace with randomly positioned objects in 3D. We demonstrate that our active inference agent is able to balance epistemic foraging and goal-driven behavior, and quantitatively outperforms both supervised and reinforcement learning baselines by more than a factor of two in terms of success rate.