Color illusions also deceive CNNs for low-level vision tasks: Analysis and implications

Color illusions also deceive CNNs for low-level vision tasks: Analysis and implications
复制标题

DOI:
10.1016/j.visres.2020.07.010
复制
发表时间:
2020-11-01
期刊:
影响因子:
1.8
通讯作者:
Malo, J.
Malo, J.
中科院分区:
心理学3区
文献类型:
--
作者:
Gomez-Villa, A.;Martin, A.;Malo, J.

文献摘要

被引文献

相似文献

视觉错觉的研究已被证明是视觉科学中一种非常有用的方法。在这项工作中,我们首先表明,虽然为自然图像中的低级别视觉任务训练的卷积神经网络(CNN)可能会受到亮度和颜色错觉的欺骗,但一些网络错觉可能与人类的感知不一致。接下来,我们分析这些相似和不同之处可能来自哪里。一方面,提出的线性特征分析解释了总体上的相似性:在为去噪或去模糊等任务训练的简单CNN中,网络的线性版本具有中心环绕的接受区,并且全局传递函数非常类似于人类在类似于人类的对手颜色空间中的非彩色和彩色对比度敏感性函数。这些相似性与长期以来的假设是一致的,该假说认为低水平视错觉是对自然环境进行优化的副产品。具体地说,这里类似人类的特征是从误差最小化中产生的。另一方面,观察到的差异一定是由于人类视觉系统的行为不能用线性近似来解释。然而,我们的研究也表明,更灵活的网络结构,具有更多的层和更高的非线性程度,实际上可能会有更差的再现视觉错觉的能力。这意味着,与视觉科学文献中的其他工作一样,关于使用CNN来研究人类视觉的警告:除了L+NL人工网络公式对视觉建模的内在局限性之外,柔性建筑的非线性行为可能很容易与视觉系统的非线性行为显著不同。
The study of visual illusions has proven to be a very useful approach in vision science. In this work we start by showing that, while convolutional neural networks (CNNs) trained for low-level visual tasks in natural images may be deceived by brightness and color illusions, some network illusions can be inconsistent with the perception of humans. Next, we analyze where these similarities and differences may come from. On one hand, the proposed linear eigenanalysis explains the overall similarities: in simple CNNs trained for tasks like denoising or deblurring, the linear version of the network has center-surround receptive fields, and global transfer functions are very similar to the human achromatic and chromatic contrast sensitivity functions in human-like opponent color spaces. These similarities are consistent with the long-standing hypothesis that considers low-level visual illusions as a by-product of the optimization to natural environments. Specifically, here human-like features emerge from error minimization. On the other hand, the observed differences must be due to the behavior of the human visual system not explained by the linear approximation. However, our study also shows that more 'flexible' network architectures, with more layers and a higher degree of nonlinearity, may actually have a worse capability of reproducing visual illusions. This implies, in line with other works in the vision science literature, a word of caution on using CNNs to study human vision: on top of the intrinsic limitations of the L + NL formulation of artificial networks to model vision, the nonlinear behavior of flexible architectures may easily be markedly different from that of the visual system.