Convolutional neural networks explain tuning properties of anterior, but not middle, face-processing areas in macaque inferotemporal cortex

Convolutional neural networks explain tuning properties of anterior, but not middle, face-processing areas in macaque inferotemporal cortex
复制标题

DOI:
10.1038/s42003-020-0945-x
复制
发表时间:
2020-05-08
影响因子:
5.9
通讯作者:
Hosoya, Haruo
Hosoya, Haruo
中科院分区:
生物学2区
文献类型:
--
作者:
Raman, Rajani;Hosoya, Haruo

文献摘要

被引文献

相似文献

Rajani Raman 和 Haruo Hosoya 研究了卷积神经网络和猕猴视觉皮层面部处理区域之间的逐层相似性。他们表明,较高的模型层表现出与前部区域类似的调整特性,例如大小和视图方面的强不变性,而中间区域更难以建模。最近的计算研究强调了卷积神经网络(CNN)和灵长类视觉腹侧流之间的逐层定量相似性。然而,这种相似性是否适用于面部选择区域(高级视觉皮层的一个子系统)尚不清楚。在这里,我们广泛研究 CNN 是否表现出之前在不同猕猴面部区域观察到的调节特性。在模拟过去对各种 CNN 模型的四次实验时,我们寻找能够定量匹配每个面部区域的多个调整属性的模型层。我们的结果表明,较高的模型层可以很好地解释前部区域的属性,而没有任何层可以同时解释中间区域的属性,并且在整个模型变化中保持一致。因此,CNN 和灵长类面部处理系统在近目标表示方面可能存在一些相似性,但在中间阶段则不太明显,因此需要替代建模,例如非分层对应或不同的计算原理。
Rajani Raman and Haruo Hosoya examined layer-wise similarities between convolutional neural networks and the face-processing areas of the macaque visual cortex. They show that higher model layers exhibit similar tuning properties of anterior areas, such as strong invariance properties in size and view, while intermediate areas were more difficult to model.Recent computational studies have emphasized layer-wise quantitative similarity between convolutional neural networks (CNNs) and the primate visual ventral stream. However, whether such similarity holds for the face-selective areas, a subsystem of the higher visual cortex, is not clear. Here, we extensively investigate whether CNNs exhibit tuning properties as previously observed in different macaque face areas. While simulating four past experiments on a variety of CNN models, we sought for the model layer that quantitatively matches the multiple tuning properties of each face area. Our results show that higher model layers explain reasonably well the properties of anterior areas, while no layer simultaneously explains the properties of middle areas, consistently across the model variation. Thus, some similarity may exist between CNNs and the primate face-processing system in the near-goal representation, but much less clearly in the intermediate stages, thus requiring alternative modeling such as non-layer-wise correspondence or different computational principles.