Towards Shape-regularized Learning for Mitigating Texture Bias in CNNs

Towards Shape-regularized Learning for Mitigating Texture Bias in CNNs
复制标题

DOI:
10.1145/3591106.3592231
复制
发表时间:
2023-06
期刊:
Proceedings of the 2023 ACM International Conference on Multimedia Retrieval
影响因子:
--
通讯作者:
Harsh Sinha;Adriana Kovashka
Harsh Sinha;Adriana Kovashka
中科院分区:
其他
文献类型:
--
作者:
Harsh Sinha;Adriana Kovashka

文献摘要

被引文献

相似文献

CNN 已成为强大的目标识别技术。然而,CNN 的测试性能取决于与训练分布的相似性。现有方法侧重于数据增强以解决域外泛化问题。相反,我们通过鼓励模型学习与从对象形状中学到的特征相关的特征来强制形状偏差。我们表明,显式的形状线索使 CNN 能够学习对看不见的图像操作具有鲁棒性的特征,即具有相同语义内容的新颖纹理。我们的模型在 Toys4K 数据集上进行了验证,该数据集包含 4179 个 3D 对象和图像对。为了量化纹理偏差,我们合成了称为 Style(GAN 的风格转移)、CueConflict(冲突纹理和语义)和 Scrambled 数据集(通过扰乱像素块来混淆语义)的数据集变体。我们的实验表明,使用形状的好处不受点云等特定形状表示的影响,而是可以从距离变换等更简单的表示中获得相同的好处。
CNNs have emerged as powerful techniques for object recognition. However, the test performance of CNNs is contingent on the similarity to training distribution. Existing methods focus on data augmentation to address out-of-domain generalization. In contrast, we enforce a shape bias by encouraging our model to learn features that correlate with those learned from the shape of the object. We show that explicit shape cues enable CNNs to learn features that are robust to unseen image manipulations i.e. novel textures with the same semantic content. Our models are validated on Toys4K dataset which consists of 4179 3D objects and image pairs. To quantify texture bias, we synthesize dataset variants called Style (style-transfer with GANs), CueConflict (conflicting texture & semantics), and Scrambled datasets (obfuscating semantics by scrambling pixel blocks). Our experiments show that the benefits of using shape is not subject to specific shape representations like point clouds, rather the same benefits can be obtained from a simpler representation such as the distance transform.