Comparing deep neural networks against humans: object recognition when the signal gets weaker

Comparing deep neural networks against humans: object recognition when the signal gets weaker
复制标题

DOI:
--
复制
发表时间:
2017-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Robert Geirhos;David H. J. Janssen;Heiko H. Schütt;Jonas Rauber;M. Bethge;Felix Wichmann
Robert Geirhos;David H. J. Janssen;Heiko H. Schütt;Jonas Rauber;M. Bethge;Felix Wichmann
中科院分区:
其他
文献类型:
--
作者:
Robert Geirhos;David H. J. Janssen;Heiko H. Schütt;Jonas Rauber;M. Bethge;Felix Wichmann

文献摘要

被引文献

相似文献

人类视觉对象识别通常是快速的,似乎毫不费力,并且在很大程度上独立于视点和对象方向。直到最近,动画视觉系统是唯一能够实现这一非凡计算壮举的系统。随着一种被称为深度神经网络(dnn)的计算机视觉算法的兴起,这种情况发生了变化,这种算法在物体识别任务上实现了人类级别的分类性能。此外,越来越多的研究报告了dnn和人类视觉系统处理物体的方式的相似性,这表明当前的dnn可能是人类视觉物体识别的良好模型。然而,最先进的深度神经网络与灵长类视觉系统之间显然存在重要的结构和处理差异。这些差异的潜在行为后果还没有得到很好的理解。我们的目标是通过比较人类和深度神经网络对图像退化的泛化能力来解决这个问题。我们发现,人类的视觉系统对图像处理,如对比度降低,加性噪声或新的图像失真,具有更强的鲁棒性。此外,当信号变弱时,我们发现人类和dnn之间的分类错误模式逐渐分化,这表明人类和当前dnn执行视觉对象识别的方式可能仍然存在显着差异。我们设想,我们的发现以及我们精心测量和免费提供的行为数据集为计算机视觉界提供了一个新的有用的基准,以提高dnn的鲁棒性,并激励神经科学家在大脑中寻找可以促进这种鲁棒性的机制。
Human visual object recognition is typically rapid and seemingly effortless, as well as largely independent of viewpoint and object orientation. Until very recently, animate visual systems were the only ones capable of this remarkable computational feat. This has changed with the rise of a class of computer vision algorithms called deep neural networks (DNNs) that achieve human-level classification performance on object recognition tasks. Furthermore, a growing number of studies report similarities in the way DNNs and the human visual system process objects, suggesting that current DNNs may be good models of human visual object recognition. Yet there clearly exist important architectural and processing differences between state-of-the-art DNNs and the primate visual system. The potential behavioural consequences of these differences are not well understood. We aim to address this issue by comparing human and DNN generalisation abilities towards image degradations. We find the human visual system to be more robust to image manipulations like contrast reduction, additive noise or novel eidolon-distortions. In addition, we find progressively diverging classification error-patterns between humans and DNNs when the signal gets weaker, indicating that there may still be marked differences in the way humans and current DNNs perform visual object recognition. We envision that our findings as well as our carefully measured and freely available behavioural datasets provide a new useful benchmark for the computer vision community to improve the robustness of DNNs and a motivation for neuroscientists to search for mechanisms in the brain that could facilitate this robustness.