Where to Look Next? Combining Static and Dynamic Proto-objects in a TVA-based Model of Visual Attention

Where to Look Next? Combining Static and Dynamic Proto-objects in a TVA-based Model of Visual Attention
复制标题

DOI:
10.1007/s12559-010-9080-1
复制
发表时间:
2010-11
影响因子:
5.4
通讯作者:
Marco Wischnewski;Anna Belardinelli;W. Schneider;Jochen J. Steil
Marco Wischnewski;Anna Belardinelli;W. Schneider;Jochen J. Steil
中科院分区:
计算机科学2区
文献类型:
--
作者:
Marco Wischnewski;Anna Belardinelli;W. Schneider;Jochen J. Steil

文献摘要

被引文献

相似文献

决定“下一步看哪里”是人类、动物和机器人注意力系统的核心功能。注意力的控制取决于三个因素,即环境的低级静态和动态视觉特征(自下而上),原型对象的中级视觉特征和任务(自上而下)。我们提出了一个新的综合计算模型,其中包括所有这些因素在一个连贯的架构基于发现和灵长类视觉系统的限制。该模型结合了静态特征的空间非均匀处理、时空运动特征和任务依赖的优先级控制,以“视觉注意理论”(TVA,[7])规定的显著性计算的首次计算实现的形式。重要的是,静态和动态处理流在视觉原型对象(即椭球形视觉单元,具有主轴的位置、大小、形状和方向等额外的中级特征)级别上融合。原型对象作为TVA过程的输入,TVA过程结合了自顶向下和自底向上的信息,用于计算注意力优先级,从而可以实现相对复杂的搜索任务。为此,分别计算的静态和动态原型对象被过滤,然后合并成一个原型对象组合图。对于每个原型对象,根据TVA以注意权重的形式计算注意优先级。下一个扫视的目标是根据任务权重最高的原型物体的重心。我们通过将其应用于几个真实世界的图像序列来说明该方法,并表明它对参数变化具有鲁棒性。
To decide “Where to look next ?” is a central function of the attention system of humans, animals and robots. Control of attention depends on three factors, that is, low-level static and dynamic visual features of the environment (bottom-up), medium-level visual features of proto-objects and the task (top-down). We present a novel integrated computational model that includes all these factors in a coherent architecture based on findings and constraints from the primate visual system. The model combines spatially inhomogeneous processing of static features, spatio-temporal motion features and task-dependent priority control in the form of the first computational implementation of saliency computation as specified by the “Theory of Visual Attention” (TVA, [7]). Importantly, static and dynamic processing streams are fused at the level of visual proto-objects, that is, ellipsoidal visual units that have the additional medium-level features of position, size, shape and orientation of the principal axis. Proto-objects serve as input to the TVA process that combines top-down and bottom-up information for computing attentional priorities so that relatively complex search tasks can be implemented. To this end, separately computed static and dynamic proto-objects are filtered and subsequently merged into one combined map of proto-objects. For each proto-object, attentional priorities in the form of attentional weights are computed according to TVA. The target of the next saccade is the center of gravity of the proto-object with the highest weight according to the task. We illustrate the approach by applying it to several real world image sequences and show that it is robust to parameter variations.