A brain-inspired object-based attention network for multiobject recognition and visual reasoning

A brain-inspired object-based attention network for multiobject recognition and visual reasoning
复制标题

DOI:
10.1101/2022.04.02.486850
复制
发表时间:
2022-04
期刊:
影响因子:
1.8
通讯作者:
Hossein Adeli;Seoyoung Ahn;G. Zelinsky
Hossein Adeli;Seoyoung Ahn;G. Zelinsky
中科院分区:
医学4区
文献类型:
--
作者:
Hossein Adeli;Seoyoung Ahn;G. Zelinsky

文献摘要

被引文献

相似文献

视觉系统使用对物体的选择性一瞥序列来支持行为目标,但是这种注意力控制是如何学习的呢?在这里,我们提出了一个编码器-解码器模型,其灵感来自于构成大脑识别-注意系统的自下而上和自上而下视觉通路的相互作用。在每次迭代中,从图像中获取新的一瞥,并通过“what”编码器(前馈、循环和胶囊层的层次结构)进行处理,以获得以对象为中心(对象文件)的表示。这种表示提供给“where”解码器,其中不断发展的循环表示提供自上而下的注意力调制,以计划随后的一瞥和影响编码器中的路由。我们展示了注意机制如何显著提高分类高度重叠数字的准确性。在需要比较两个物体的视觉推理任务中,我们的模型达到了近乎完美的准确性,并且在推广到看不见的刺激方面显着优于大型模型。我们的工作证明了基于对象的注意机制对对象的顺序瞥见的好处。
The visual system uses sequences of selective glimpses to objects to support behavioral goals, but how is this attention control learned? Here we present an encoder-decoder model inspired by the interacting bottom-up and top-down visual pathways making up the recognition-attention system in the brain. At every iteration, a new glimpse is taken from the image and is processed through the ‘what’ encoder, a hierarchy of feedforward, recurrent, and capsule layers, to obtain an object-centric (object-file) representation. This representation feeds to the ‘where’ decoder, where the evolving recurrent representation provides top-down attentional modulation to plan subsequent glimpses and impact routing in the encoder. We demonstrate how the attention mechanism significantly improves the accuracy of classifying highly overlapping digits. In a visual reasoning task requiring comparison of two objects, our model achieves near-perfect accuracy and significantly outperforms larger models in generalizing to unseen stimuli. Our work demonstrates the benefits of object-based attention mechanisms taking sequential glimpses of objects.