Disentangling bottom-up versus top-down and low-level versus high-level influences on eye movements over time

Disentangling bottom-up versus top-down and low-level versus high-level influences on eye movements over time
复制标题

DOI:
10.1167/19.3.1
复制
发表时间:
2019-03-01
期刊:
影响因子:
1.8
通讯作者:
Wichmann, Felix A.
Wichmann, Felix A.
中科院分区:
医学4区
文献类型:
--
作者:
Schuett, Heiko H.;Rothkegel, Lars O. M.;Wichmann, Felix A.

文献摘要

被引文献

相似文献

自下而上,自上而下以及低级和高级因素会影响我们在观看自然场景时固定的位置。但是,每个因素的重要性以及它们如何相互作用仍然是辩论的问题。在这里,我们通过分析它们随时间的影响来解散这些因素。为此,我们开发了一个显着性模型,该模型基于最新的早期空间视觉模型的内部表示,以测量低级,自下而上的因素。为了衡量高级,自下而上的特征的影响,我们使用了最近的深层神经网络显着性模型。为了解释自上而下的影响,我们在两个大型数据集上评估了具有不同任务的大型数据集的模型:首先,记忆任务,其次是搜索任务。我们的结果将视觉场景探索分为三个阶段的支持提供了支持:第一个扫视,最初的引导探索,其特征是固定密度的逐渐扩大,以及大约10个固定装置后达到的稳定状态。在初始探索和稳定状态下的扫视目标选择与相似的感兴趣领域有关,这些领域在包括高级特征时可以更好地预测。在搜索数据集中,固定位置主要由自上而下的过程确定。相反,第一个固定遵循不同的固定密度,并包含强烈的中央固定偏置。尽管如此,第一固定是由图像属性强烈指导的,并且最早在图像发作后200 ms,可以通过高级信息更好地预测固定。我们得出的结论是,任何低级,自下而上的因素主要仅限于第一个扫视的产生。当考虑高级特征时,可以更好地解释所有扫视,稍后,这种高水平的自下而上的控制可以被自上而下的影响否决。
Bottom-up and top-down as well as low-level and high-level factors influence where we fixate when viewing natural scenes. However, the importance of each of these factors and how they interact remains a matter of debate. Here, we disentangle these factors by analyzing their influence over time. For this purpose, we develop a saliency model that is based on the internal representation of a recent early spatial vision model to measure the low-level, bottom-up factor. To measure the influence of high-level, bottom-up features, we use a recent deep neural network-based saliency model. To account for top-down influences, we evaluate the models on two large data sets with different tasks: first, a memorization task and, second, a search task. Our results lend support to a separation of visual scene exploration into three phases: the first saccade, an initial guided exploration characterized by a gradual broadening of the fixation density, and a steady state that is reached after roughly 10 fixations. Saccade-target selection during the initial exploration and in the steady state is related to similar areas of interest, which are better predicted when including high-level features. In the search data set, fixation locations are determined predominantly by top-down processes. In contrast, the first fixation follows a different fixation density and contains a strong central fixation bias. Nonetheless, first fixations are guided strongly by image properties, and as early as 200 ms after image onset, fixations are better predicted by high-level information. We conclude that any low-level, bottom-up factors are mainly limited to the generation of the first saccade. All saccades are better explained when high-level features are considered, and later, this high-level, bottom-up control can be overruled by top-down influences.