Large-Scale, High-Resolution Comparison of the Core Visual Object Recognition Behavior of Humans, Monkeys, and State-of-the-Art Deep Artificial Neural Networks

Large-Scale, High-Resolution Comparison of the Core Visual Object Recognition Behavior of Humans, Monkeys, and State-of-the-Art Deep Artificial Neural Networks
复制标题

DOI:
10.1523/jneurosci.0388-18.2018
复制
发表时间:
2018-08-15
影响因子:
5.3
通讯作者:
DiCarlo, James J.
DiCarlo, James J.
中科院分区:
医学1区
文献类型:
--
作者:
Rajalingham, Rishi;Issa, Elias B.;DiCarlo, James J.

文献摘要

被引文献

相似文献

灵长类动物,包括人类,通常可以一目了然地识别视觉图像中的对象,尽管自然发生的身份保留图像转换(例如,视点的变化)。神经科学的一个主要目标是揭示神经元水平的机制模型,通过预测每一张图像的灵长类动物的表现来定量解释这种行为。在这里,我们将这一严格的行为预测测试应用于灵长类视觉的领先机械模型(具体而言,深度、卷积、人工神经网络;ANN),方法是直接将它们的行为特征与人类和猕猴的行为特征进行比较。使用人类和猴子心理物理学的高通量数据收集系统,我们收集了1472名匿名人和5只雄性猕猴的100多万次行为测试,在276个二进制对象辨别任务中收集了2400张图像。与以前的工作一致,我们观察到最先进的用于视觉分类的深度前馈卷积神经网络(称为DCNNIC模型)准确地预测了对象级别混淆的灵长类模式。然而,当我们检查每个物体辨别任务中单个图像的行为表现时,我们发现所有测试的DCNNIC模型都对灵长类动物的行为表现具有显著的不可预测性,而且这种预测失败既不是简单的图像属性所解释的,也不是通过简单的模型修改来挽救的。这些结果表明,现有的DCNNIC模型不能解释灵长类动物图像水平的行为模式,需要新的人工神经网络模型来更准确地捕捉灵长类物体视觉的神经机制。为此,大规模、高分辨率的灵长类行为基准,如这里获得的那些,可以作为发现此类模型的直接指南。
Primates, including humans, can typically recognize objects in visual images at a glance despite naturally occurring identity-preserving image transformations (e.g., changes in viewpoint). A primary neuroscience goal is to uncover neuron-level mechanistic models that quantitatively explain this behavior by predicting primate performance for each and every image. Here, we applied this stringent behavioral prediction test to the leading mechanistic models of primate vision (specifically, deep, convolutional, artificial neural networks; ANNs) by directly comparing their behavioral signatures against those of humans and rhesus macaque monkeys. Using high-throughput data collection systems for human and monkey psychophysics, we collected more than one million behavioral trials from 1472 anonymous humans and five male macaque monkeys for 2400 images over 276 binary object discrimination tasks. Consistent with previous work, we observed that state-of-the-art deep, feedforward convolutional ANNs trained for visual categorization (termed DCNNIC models) accurately predicted primate patterns of object-level confusion. However, when we examined behavioral performance for individual images within each object discrimination task, we found that all tested DCNNIC models were significantly nonpredictive of primate performance and that this prediction failure was not accounted for by simple image attributes nor rescued by simple model modifications. These results show that current DCNNIC models cannot account for the image-level behavioral patterns of primates and that new ANN models are needed to more precisely capture the neural mechanisms underlying primate object vision. To this end, large-scale, high-resolution primate behavioral benchmarks such as those obtained here could serve as direct guides for discovering such models.