Hiding a plane with a pixel: examining shape-bias in CNNs and the benefit of building in biological constraints

Hiding a plane with a pixel: examining shape-bias in CNNs and the benefit of building in biological constraints
复制标题

DOI:
10.1016/j.visres.2020.04.013
复制
发表时间:
2020-09-01
期刊:
影响因子:
1.8
通讯作者:
Bowers, Jeffrey S.
Bowers, Jeffrey S.
中科院分区:
心理学3区
文献类型:
--
作者:
Malhotra, Gaurav;Evans, Benjamin D.;Bowers, Jeffrey S.

文献摘要

被引文献

相似文献

当深度卷积神经网络(cnn)在原始数据上“端到端”训练时,它们在早期层中开发的一些特征检测器类似于早期视觉皮层中发现的表征。这一结果被用来在深度学习系统和人类视觉感知之间建立相似之处。在这项研究中,我们表明,当cnn被端到端训练时,它们会学习基于数据集中预测类别的任何特征对图像进行分类。这可能导致奇怪的结果,cnn学习特殊的特征,如高频噪声样掩模。在极端情况下,我们的结果展示了基于单个像素的图像分类。这些特征在人类物体识别中不太可能发挥任何作用,因为实验一再表明人类对形状有强烈的偏好。通过对标准高性能cnn的一系列实证研究,我们表明这些网络不会仅仅通过正则化方法或更生态合理的训练制度来发展形状偏差。这些结果对简单地在标准cnn中端到端学习导致出现与人类视觉系统相似的表征的假设提出了质疑。在论文的第二部分,我们表明,当我们放弃端到端学习并引入硬连线Gabor滤波器来模仿V1的早期视觉处理时,cnn对这些特殊特征的依赖程度降低了。
When deep convolutional neural networks (CNNs) are trained "end-to-end" on raw data, some of the feature detectors they develop in their early layers resemble the representations found in early visual cortex. This result has been used to draw parallels between deep learning systems and human visual perception. In this study, we show that when CNNs are trained end-to-end they learn to classify images based on whatever feature is predictive of a category within the dataset. This can lead to bizarre results where CNNs learn idiosyncratic features such as high-frequency noise-like masks. In the extreme case, our results demonstrate image categorisation on the basis of a single pixel. Such features are extremely unlikely to play any role in human object recognition, where experiments have repeatedly shown a strong preference for shape. Through a series of empirical studies with standard high-performance CNNs, we show that these networks do not develop a shape-bias merely through regularisation methods or more ecologically plausible training regimes. These results raise doubts over the assumption that simply learning end-to-end in standard CNNs leads to the emergence of similar representations to the human visual system. In the second part of the paper, we show that CNNs are less reliant on these idiosyncratic features when we forgo end-to-end learning and introduce hard-wired Gabor filters designed to mimic early visual processing in V1.