Orthogonal Representations of Object Shape and Category in Deep Convolutional Neural Networks and Human Visual Cortex

Orthogonal Representations of Object Shape and Category in Deep Convolutional Neural Networks and Human Visual Cortex
复制标题

DOI:
10.1038/s41598-020-59175-0
复制
发表时间:
2020-02-12
期刊:
影响因子:
4.6
通讯作者:
Op de Beeck, Hans
Op de Beeck, Hans
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Zeman, Astrid A.;Ritchie, J. Brendan;Op de Beeck, Hans

文献摘要

被引文献

相似文献

深度卷积神经网络(cnn)作为视觉物体识别的基准模型越来越受关注,其性能现在已经超过了人类。虽然cnn可以准确地将一张图像分配到潜在的数千个类别中,但网络性能可能是调整为表示对象的视觉形状而不是对象类别的层的结果,因为两者在自然图像中经常混淆。使用两个明确分离形状和类别的刺激集,我们将这两种类型的信息与多个cnn的每一层相关联。通过将人工表征与神经表征相关联,我们还比较了CNN输出与沿人类视觉腹侧流的fMRI激活。我们发现CNN编码类别信息独立于形状,在所有测试的CNN架构中,在最终的完全连接层达到峰值。将cnn与fMRI脑数据进行比较,发现早期视觉皮层(V1)和cnn的早期层编码形状信息。颞叶前部腹侧皮层编码类别信息,与cnn的最后一层相关性最好。在人类视觉腹侧通路上发现的形状和类别之间的相互作用在多个深度网络中得到了回应。我们的研究结果表明,cnn表示类别信息独立于形状,很像人类的视觉系统。
Deep Convolutional Neural Networks (CNNs) are gaining traction as the benchmark model of visual object recognition, with performance now surpassing humans. While CNNs can accurately assign one image to potentially thousands of categories, network performance could be the result of layers that are tuned to represent the visual shape of objects, rather than object category, since both are often confounded in natural images. Using two stimulus sets that explicitly dissociate shape from category, we correlate these two types of information with each layer of multiple CNNs. We also compare CNN output with fMRI activation along the human visual ventral stream by correlating artificial with neural representations. We find that CNNs encode category information independently from shape, peaking at the final fully connected layer in all tested CNN architectures. Comparing CNNs with fMRI brain data, early visual cortex (V1) and early layers of CNNs encode shape information. Anterior ventral temporal cortex encodes category information, which correlates best with the final layer of CNNs. The interaction between shape and category that is found along the human visual ventral pathway is echoed in multiple deep networks. Our results suggest CNNs represent category information independently from shape, much like the human visual system.