On Human-like Biases in Convolutional Neural Networks for the Perception of Slant from Texture

On Human-like Biases in Convolutional Neural Networks for the Perception of Slant from Texture
复制标题

DOI:
10.1145/3613451
复制
发表时间:
2023-08
影响因子:
1.6
通讯作者:
Yuanhao Wang;Qian Zhang;Celine Aubuchon;Jovan T. Kemp;F. Domini;J. Tompkin
Yuanhao Wang;Qian Zhang;Celine Aubuchon;Jovan T. Kemp;F. Domini;J. Tompkin
中科院分区:
计算机科学4区
文献类型:
--
作者:
Yuanhao Wang;Qian Zhang;Celine Aubuchon;Jovan T. Kemp;F. Domini;J. Tompkin

文献摘要

相似文献

深度估计是3D感知的基础,众所周知,人类对深度的估计是有偏见的。研究了在不同的观察条件(视场)和表面参数(倾斜和纹理不规则性)下,卷积神经网络(CNN)在预测纹理表面的曲率符号和深度时是否存在偏差。这一假设源于这样一个想法,即由局部邻域描述的纹理梯度--这是人类视觉文献中确定的线索--也可以在卷积神经网络中表示。为此,我们对具有随机Polka点图案的倾斜表面的渲染训练了无监督和有监督的CNN模型,并分析了它们的内部潜在表示。结果表明,在所有实验中,无监督模型与人类有类似的预测偏差,而有监督的CNN模型没有表现出类似的偏差。无监督模型的隐含空间可以线性地分解为表示视场和光学倾斜的轴。对于受监督的模型,这种能力因模型体系结构和监督的类型(连续倾斜与倾斜迹象)而有很大不同。尽管这项研究没有提到任何共同的机制,但这些发现表明,无监督的CNN模型可以分享与人类视觉系统相似的预测。代码:githorb.com/brownvc/slant-cnn-biases。
Depth estimation is fundamental to 3D perception, and humans are known to have biased estimates of depth. This study investigates whether convolutional neural networks (CNNs) can be biased when predicting the sign of curvature and depth of surfaces of textured surfaces under different viewing conditions (field of view) and surface parameters (slant and texture irregularity). This hypothesis is drawn from the idea that texture gradients described by local neighborhoods—a cue identified in human vision literature—are also representable within convolutional neural networks. To this end, we trained both unsupervised and supervised CNN models on the renderings of slanted surfaces with random Polka dot patterns and analyzed their internal latent representations. The results show that the unsupervised models have similar prediction biases as humans across all experiments, while supervised CNN models do not exhibit similar biases. The latent spaces of the unsupervised models can be linearly separated into axes representing field of view and optical slant. For supervised models, this ability varies substantially with model architecture and the kind of supervision (continuous slant vs. sign of slant). Even though this study says nothing of any shared mechanism, these findings suggest that unsupervised CNN models can share similar predictions to the human visual system. Code: github.com/brownvc/Slant-CNN-Biases.