Two Visual Systems and Their Eye Movements: Evidence from Static and Dynamic Scene Perception

Two Visual Systems and Their Eye Movements: Evidence from Static and Dynamic Scene Perception
复制标题

DOI:
--
复制
发表时间:
2005
期刊:
--
影响因子:
--
通讯作者:
B. Velichkovsky;M. Joos;J. Helmert;S. Pannasch
B. Velichkovsky;M. Joos;J. Helmert;S. Pannasch
中科院分区:
其他
文献类型:
--
作者:
B. Velichkovsky;M. Joos;J. Helmert;S. Pannasch

文献摘要

被引文献

相似文献

两种视觉系统及其眼动:来自静态和动态场景知觉的证据。放大图片作者:Markus Joos,Jens R. Helmert和塞巴斯蒂安Pannasch(velich {joosy helmert pannasch}@psychomail.tu-dresden.de)德累斯顿工业大学应用认知研究/心理学III Mommsenstrasse 13 Dresden,D-01062德国摘要灵长类动物大脑中两种不同视觉通路的存在是进化、神经生理学、运动控制和神经心理学研究的一个持久主题。作为认知神经科学中最广泛引用的结果之一,这种区别在不同的伪装下(例如环境与焦点或视觉与认知)经历了几十年的批判性分析。然而,这两个处理流之间的相互作用,在解决日常任务仍然是一个悬而未决的问题。特别是,它们如何引导眼球运动,它们最直接的输出?我们最近在模拟驾驶环境中对危险感知的研究结果表明,眼球运动参数的特定组合表明这两个系统中的任何一个参与。在进一步的实验中,我们试图通过测试与这两种模式相关的记忆表征的假设来验证这些参数。在对各种真实的世界场景进行简短的介绍后,受试者必须识别出其中的剪切部分,这些剪切部分是根据他们的固定参数选择的。还提供了从未见过的图片(捕获试验)的随机剪切。结果证实了我们的假设:切断对应于大概的焦点模式的处理更好地认识到比切断同样固定在周围的探索过程中。关键词:主动视觉;背侧和背侧流;周围和焦点注意;场景感知;识别;眼动。在当代视觉认知的研究中,人们可以发现几个研究集群,它们之间只有松散的联系(类似的论点,见Simons & Rensink,2005)。例如,一个尽管仍有争议的激烈讨论是,在观看场景时,视觉信息如何甚至是否在扫视中被保留(Bridgeman,货车der Heijden和Velichkovsky,欧文的跨扫视记忆的对象文件理论强调了视觉注意力在场景中的局部视觉信息被表示或不被表示中的关键作用(欧文,1992)。在视觉场景中关注一个对象可以将其特征绑定到一个统一的对象描述中(Treisman,1988)。这个对象描述被链接到位置主地图中的空间位置,形成视觉短期记忆(VSTM)中的临时表示。根据欧文(1992)的观点,在VSTM中一次可以容纳三到四个离散的物体。最后,目标文件是跨扫视记忆的主要内容,提供从一个注视到下一个注视的局部连续性。最近,相干理论(Rensink,2000 a,2000 b)提出了一个类似的解释,尽管视觉信息获取的“快照”特征,我们周围的世界是稳定的,连贯的,丰富的细节。就像在客体文件理论中一样,视觉注意是将感觉特征绑定到连贯的客体表征中的前提,这可以在VSTM中保持,以防止它们被破坏,如扫视。在集中注意力之前,只有有限的临时和空间一致性的原型对象在整个视野中平行形成,但在其位置出现新刺激时会发生变化并被替换。在注意力撤回时,连贯的对象再次分解成其组成的原型对象,没有或只有很少的注意力后效(Rensink,2000 a)。这两种视觉瞬变假设(Hollingworth &亨德森,2002)都依赖于这样一种想法,即视觉场景的完整度量表示既不可能也不必要。O 'Regan和Noe(2001)甚至更进一步地说(实际上是对布鲁克斯,1991,第139页的释义)“世界作为它自己的记忆”。然而,这些场景感知方法的一个缺点是,它们将眼睛运动视为没有进一步兴趣的机械事件。该分类简单地基于测量受试者是否将他们的眼睛稳定地保持在指定位置(注视)或正在进行急动的眼睛运动(扫视)。在大多数情况下,视觉刺激(例如,场景或对象)的呈现或消失是相对于扫视的开始或结束来执行的。潜在的加工,这可能会反映在固定时间的变化,被忽略。还有另一种研究分析了视觉注视的持续时间,包括任务的复杂性、处理水平或技能,特别是阅读任务(Velichkovsky,1999)。在Unema,Pannasch,Joos和Velichkovsky(2005)最近的一项研究中,受试者观看计算机生成的包含不同内部的房间图像,以便能够回答有关房间内物体分布或特定物体存在/不存在的问题。作者发现,在整个任务和观看时间内,注视持续时间和扫视幅度的比率有明显的变化。在图像检查开始时,
Two Visual Systems and their Eye Movements: Evidence from Static and Dynamic Scene Perception Boris M. Velichkovsky, Markus Joos, Jens R. Helmert, and Sebastian Pannasch (velich {joosy helmert pannasch}@psychomail.tu-dresden.de) Dresden University of Technology Applied Cognitive Research/ Psychology III Mommsenstrasse 13 Dresden, D-01062 Germany Abstract The existence of two distinct visual pathways in the primate brain is a persistent theme for evolutionary, neurophysio- logical, motor control and neuropsychological research. As one of the most widely cited results in cognitive neuroscience, this distinction has survived decades of critical analysis under different guises (e.g. ambient vs. focal or visuomotor vs. cognitive). However, the interplay between these two processing streams in the solution of everyday tasks remains to be an unresolved issue. In particular, how do they guide eye movements, their most immediate output? Results from our recent study on hazard perception in a simulated driving environment demonstrated that specific combinations of eye movement parameters are indicative to an involvement of either of the two systems. In a further experiment, we tried to validate these parameters by testing assumptions about memory representations related to these two modes. After a short presentation of various real world scenes, subjects had to recognize cut-outs from them, which were selected according to their fixation parameters. Random cut-outs from not seen pictures (catch trials) were also presented. The results confirmed our hypothesis: cut-outs corresponding to presumably focal mode of processing were better recognized than cut-outs similarly fixated in the course of ambient exploration. Keywords: Active Vision; Dorsal and Ventral Streams; Ambient and Focal Attention; Scene Perception; Recognition; Eye Movements. Introduction In contemporary studies of visual cognition, one can discover several clusters of research that are only loosely connected to each other (for similar arguments, see Simons & Rensink, 2005). An intensive albeit still controversial discussion is, for instance, how and even whether visual information is retained across saccades while viewing a scene (Bridgeman, Van der Heijden, and Velichkovsky, Irwin's object file theory of transsaccadic memory emphasises the crucial role of visual attention in what local visual information from a scene is or is not represented (Irwin, 1992). Attending an object in a visual scene allows binding its features into a unified object description (Treisman, 1988). This object description is linked to a spatial position in a master map of locations, forming a temporary representation in visual short-term memory ( VSTM ). According to Irwin (1992) three to four discrete objects can be hold at a time in VSTM . Finally, object files are the primary content of transsaccadic memory providing local continuity from one fixation to the next. More recently, coherence theory (Rensink, 2000a, 2000b) proposed a similar explanation of the fact that despite the 'snapshot-like' character of visual information acquisition the world around us is experienced as being stable, coherent and richly detailed. Just as in object file theory visual attention is the premise to bind sensory features into a coherent object representation, which can be hold in VSTM preventing them from disruptions like saccades. Prior to focused attention, proto-objects with only limited temporary and spatial coherence are formed in parallel across the visual field but being volatile and replaced on appearance of a new stimulus at their position. On withdrawal of attention a coherent object resolves into its constituent proto-objects again leaving no or only little after-effect of attention (Rensink, 2000a). Both these visual transience hypotheses (Hollingworth & Henderson, 2002) rely on the idea, that a complete metric representation of a visual scene is neither possible nor necessary. O'Regan and Noe (2001) go even further saying (actually paraphrasing Brooks, 1991, p. 139) “the world serves as its own memory”. One drawback of these approaches to scene perception is, however, that they consider eye movements as mechanical events of no further interest. The classification is simply based on measuring whether the subject is holding their eyes stable at a designated position (fixation) or is doing a jerky eye movement (saccade). Presentation or extinction of visual stimuli (e.g. a scene or an object) is in the majority of cases executed in relation to the start or end of a saccade. The underlying processing, which may be reflected in the variation of fixation durations, is neglected. There is another line of research analyzing the duration of visual fixations in terms of task complexity, levels of processing or skills, especially for reading tasks (Velichkovsky, 1999). In a recent study by Unema, Pannasch, Joos, and Velichkovsky (2005) subjects viewed computer generated images of rooms containing different interior, in order to be able to answer questions about the distribution of objects within the room or about the presence/absence of particular objects. The authors found a clear shift of the ratio of fixation durations and saccadic amplitudes across the tasks and also over the viewing time. At the beginning of image inspection fixations with shorter durations and saccades with longer amplitudes were