Can machine learning account for human visual object shape similarity judgments?

Can machine learning account for human visual object shape similarity judgments?
复制标题

DOI:
10.1016/j.visres.2019.12.001
复制
发表时间:
2020-01
期刊:
影响因子:
1.8
通讯作者:
Joseph German;R. Jacobs
Joseph German;R. Jacobs
中科院分区:
心理学3区
文献类型:
--
作者:
Joseph German;R. Jacobs

文献摘要

被引文献

相似文献

我们描述和分析了度量学习系统的性能,包括深度神经网络(DNN),在一个新的数据集上的人类视觉对象形状相似性判断的自然,部分为基础的对象被称为“Fribbles”。与以前的研究,要求参与者判断相似性时,从一个单一的观点呈现的对象或场景,我们从多个角度呈现Fribbles,并要求参与者判断形状相似性的视点不变的方式。使用基于像素或基于DNN的表示进行训练无法解释我们的实验数据,但是使用视点不变的基于部件的表示进行训练的度量产生了很好的拟合。我们还发现,虽然神经网络可以学习提取基于部分的表示,因此应该能够学习模型我们的数据网络训练与“三重损失”功能的基础上相似性判断没有表现得很好。我们分析了这种失败,提供了一个数学描述的度量学习目标函数和三重损失函数之间的关系。神经网络的性能差似乎是由于网络权值空间的优化问题的非凸性。我们的结论是,视点不敏感是人类视觉形状感知的一个重要方面,神经网络和其他机器学习方法将需要学习视点不敏感的表示,以考虑人们的视觉对象形状相似性判断。
We describe and analyze the performance of metric learning systems, including deep neural networks (DNNs), on a new dataset of human visual object shape similarity judgments of naturalistic, part-based objects known as “Fribbles”. In contrast to previous studies which asked participants to judge similarity when objects or scenes were rendered from a single viewpoint, we rendered Fribbles from multiple viewpoints and asked participants to judge shape similarity in a viewpoint-invariant manner. Metrics trained using pixel-based or DNN-based representations fail to explain our experimental data, but a metric trained with a viewpoint-invariant, part-based representation produces a good fit. We also find that although neural networks can learn to extract the part-based representation—and therefore should be capable of learning to model our data—networks trained with a “triplet loss” function based on similarity judgments do not perform well. We analyze this failure, providing a mathematical description of the relationship between the metric learning objective function and the triplet loss function. The poor performance of neural networks appears to be due to the nonconvexity of the optimization problem in network weight space. We conclude that viewpoint insensitivity is a critical aspect of human visual shape perception, and that neural network and other machine learning methods will need to learn viewpoint-insensitive representations in order to account for people’s visual object shape similarity judgments.