Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex

Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex
复制标题

DOI:
10.48550/arxiv.2306.03779
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Drew Linsley;I. F. Rodriguez;Thomas Fel;Michael Arcaro;Saloni Sharma;M. Livingstone;Thomas Serre
Drew Linsley;I. F. Rodriguez;Thomas Fel;Michael Arcaro;Saloni Sharma;M. Livingstone;Thomas Serre
中科院分区:
其他
文献类型:
--
作者:
Drew Linsley;I. F. Rodriguez;Thomas Fel;Michael Arcaro;Saloni Sharma;M. Livingstone;Thomas Serre

文献摘要

相似文献

在过去十年中,计算神经科学中最具影响力的发现之一是,深度神经网络(DNN)的对象识别准确性与其预测下颞叶(IT)皮层对自然图像的神经反应的能力相关。这一发现支持了长期以来的理论,即物体识别是视觉皮层的核心目标,并表明更准确的DNN将作为IT神经元对图像反应的更好模型。从那时起,深度学习经历了一场规模革命:在数十亿张图像上训练的数十亿个参数规模的DNN在包括对象识别在内的视觉任务中与人类竞争或超越人类。今天的DNN在预测IT神经元对图像的反应方面变得更加准确,因为它们在物体识别方面变得更加准确?令人惊讶的是,在三个独立的实验中,我们发现情况并非如此。DNN已经成为越来越糟糕的IT模型,因为它们在ImageNet上的准确性有所提高。为了理解为什么DNN会经历这种权衡,并评估它们是否仍然是模拟视觉系统的合适范例,我们转向IT的记录,这些记录捕获了自然图像引起的神经元活动的空间分辨地图。这些神经元活动图显示,在ImageNet上训练的DNN学习依赖于与IT编码的视觉特征不同的视觉特征,并且随着其准确性的提高,这个问题会变得更加复杂。我们成功地解决了这个问题与神经协调器,即插即用的训练程序DNN对齐他们的学习表示与人类。我们的研究结果表明,协调的DNN打破了ImageNet准确性和神经预测准确性之间的权衡,并为更准确的生物视觉模型提供了一条道路。
One of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the long-held theory that object recognition is a core objective of the visual cortex, and suggested that more accurate DNNs would serve as better models of IT neuron responses to images. Since then, deep learning has undergone a revolution of scale: billion parameter-scale DNNs trained on billions of images are rivaling or outperforming humans at visual tasks including object recognition. Have today's DNNs become more accurate at predicting IT neuron responses to images as they have grown more accurate at object recognition? Surprisingly, across three independent experiments, we find this is not the case. DNNs have become progressively worse models of IT as their accuracy has increased on ImageNet. To understand why DNNs experience this trade-off and evaluate if they are still an appropriate paradigm for modeling the visual system, we turn to recordings of IT that capture spatially resolved maps of neuronal activity elicited by natural images. These neuronal activity maps reveal that DNNs trained on ImageNet learn to rely on different visual features than those encoded by IT and that this problem worsens as their accuracy increases. We successfully resolved this issue with the neural harmonizer, a plug-and-play training routine for DNNs that aligns their learned representations with humans. Our results suggest that harmonized DNNs break the trade-off between ImageNet accuracy and neural prediction accuracy that assails current DNNs and offer a path to more accurate models of biological vision.