Modeling the influence of task on attention

Modeling the influence of task on attention
复制标题

DOI:
10.1016/j.visres.2004.07.042
复制
发表时间:
2005-01-01
期刊:
影响因子:
1.8
通讯作者:
Itti, L
Itti, L
中科院分区:
心理学3区
文献类型:
--
作者:
Navalpakkam, V;Itti, L

文献摘要

被引文献

相似文献

我们提出了一个在现实世界场景中特定任务的视觉注意引导的计算模型。我们的模型强调了生物视觉中重要的四个方面:确定实体的任务相关性,将注意力偏向于期望目标的低层次视觉特征,使用相同的低层次特征识别这些目标,以及在每个场景位置逐步构建任务相关性的视觉地图。在给定以关键词形式的任务定义后,该模型首先利用存储在长期记忆中的先验知识,确定任务相关实体并将其存储在工作记忆中。它试图通过将其视觉注意系统与实体学习的低级特征相偏差来检测最相关的实体。它关注场景中最突出的位置。并尝试通过与存储在长期记忆中的对象表示进行分层匹配来识别被关注的对象。它用被识别实体的任务相关性更新工作记忆,用被识别实体的位置和相关性更新任务相关性地形图。该模型在三种类型的任务上进行了测试:在343张自然和合成图像中进行单目标检测,其中目标的偏置平均加速目标检测两倍以上;28幅自然图像的序列多目标检测,其中偏置、识别、工作记忆和长期记忆有助于快速找到所有目标;并从高速公路上行驶时拍摄的视频片段中学习汽车可能位置的地图。该模型在单个特征和特征连词搜索方面的表现与已有的心理物理数据一致。我们的生物驱动架构的这些结果表明,该模型可能为涉及复杂任务驱动的视觉行为的许多大脑过程提供合理的近似。(C) 2004 Elsevier Ltd.版权所有。
We propose a computational model for the task-specific guidance of visual attention in real-world scenes. Our model emphasizes four aspects that are important in biological vision: determining task-relevance of an entity, biasing attention for the low-level visual features of desired targets, recognizing these targets using the same low-level features, and incrementally building a visual map of task-relevance at every scene location. Given a task definition in the form of keywords, the model first determines and stores the task-relevant entities in working memory, using prior knowledge stored in long-term memory. It attempts to detect the most relevant entity by biasing its visual attention system with the entity's learned low-level features. It attends to the most salient location in the scene. and attempts to recognize the attended object through hierarchical matching against object representations stored in long-term memory. It updates its working memory with the task-relevance of the recognized entity and updates a topographic task-relevance map with the location and relevance of the recognized entity. The model is tested on three types of tasks: single-target detection in 343 natural and synthetic images, where biasing for the target accelerates target detection over twofold on average; sequential multiple-target detection in 28 natural images, where biasing, recognition, working memory and long term memory contribute to rapidly finding all targets; and learning a map of likely locations of cars from a video clip filmed while driving on a highway. The model's performance on search for single features and feature conjunctions is consistent with existing psychophysical data. These results of our biologically-motivated architecture suggest that the model may provide a reasonable approximation to many brain processes involved in complex task-driven visual behaviors. (C) 2004 Elsevier Ltd. All rights reserved.