A fast, invariant representation for human action in the visual system

A fast, invariant representation for human action in the visual system
复制标题

DOI:
10.1152/jn.00642.2017
复制
发表时间:
2018-02-01
影响因子:
2.5
通讯作者:
Poggio, Tomaso
Poggio, Tomaso
中科院分区:
医学3区
文献类型:
--
作者:
Isik, Leyla;Tacchetti, Andrea;Poggio, Tomaso

文献摘要

被引文献

相似文献

人类可以毫不费力地在复杂的变化中识别他人的行为,比如视点的变化。一些研究已经确定了大脑中参与不变动作识别的区域;然而,潜在的神经计算仍然知之甚少。我们使用脑磁图解码和一组控制良好的、自然的五种动作(跑、走、跳、吃、喝)的视频数据集,这些视频由不同的演员在不同的视角进行,以研究用于识别复杂转换中的动作的计算步骤。特别是,我们想知道大脑何时会区分不同的动作,以及它何时会以一种不受3D视点变化影响的方式进行区分。我们测量了当受试者观看完整视频以及形式耗尽和运动耗尽刺激时,不变和非不变动作解码之间的延迟差异。在完整视频中,我们无法检测到不变和非不变动作识别之间的解码延迟或时间概况的差异。然而,当从刺激集中去除形式或运动信息时,我们观察到不变动作解码的减少和延迟。我们的研究结果表明,大脑在识别动作的同时建立对复杂变换的不变性,而形式和运动信息对于快速、不变性的动作识别至关重要。新的和值得注意的是,尽管改变了视觉外观,人类的大脑仍然可以快速识别动作。我们使用神经时序数据来揭示这种能力背后的计算。我们发现在200毫秒内的动作可以从脑磁图数据中读取出来,并且这种表示对视点的变化是不变的。我们发现这种快速动作解码需要形式和运动,这表明大脑快速整合复杂的时空特征以形成不变的动作表征。
Humans can effortlessly recognize others' actions in the presence of complex transformations, such as changes in viewpoint. Several studies have located the regions in the brain involved in invariant action recognition; however, the underlying neural computations remain poorly understood. We use magnetoencephalography decoding and a data set of well-controlled, naturalistic videos of five actions (run, walk, jump, eat, drink) performed by different actors at different viewpoints to study the computational steps used to recognize actions across complex transformations. In particular, we ask when the brain discriminates between different actions, and when it does so in a manner that is invariant to changes in 3D viewpoint. We measure the latency difference between invariant and noninvariant action decoding when subjects view full videos as well as form-depleted and motion-depleted stimuli. We were unable to detect a difference in decoding latency or temporal profile between invariant and noninvariant action recognition in full videos. However, when either form or motion information is removed from the stimulus set, we observe a decrease and delay in invariant action decoding. Our results suggest that the brain recognizes actions and builds invariance to complex transformations at the same time and that both form and motion information are crucial for fast, invariant action recognition.NEW & NOTEWORTHY The human brain can quickly recognize actions despite transformations that change their visual appearance. We use neural timing data to uncover the computations underlying this ability. We find that within 200 ms action can be read out of magnetoencephalography data and that this representation is invariant to changes in viewpoint. We find form and motion are needed for this fast action decoding, suggesting that the brain quickly integrates complex spatiotemporal features to form invariant action representations.