Deep convolutional models improve predictions of macaque V1 responses to natural images

Deep convolutional models improve predictions of macaque V1 responses to natural images
复制标题

DOI:
10.1371/journal.pcbi.1006897
复制
发表时间:
2019-04-01
影响因子:
4.3
通讯作者:
Ecker, Alexander S.
Ecker, Alexander S.
中科院分区:
生物学2区
文献类型:
--
作者:
Cadena, Santiago A.;Denfield, George H.;Ecker, Alexander S.

文献摘要

被引文献

相似文献

尽管经过几十年的努力,我们最好的初级视觉皮层(V1)模型在用自然刺激探测时仍然预测得很差,突出了我们对V1中非线性计算的有限理解。最近,出现了两种基于深度学习的方法来对这些非线性计算进行建模:从对象识别训练的人工神经网络进行转移学习,以及在大量神经元上进行端到端训练的数据驱动卷积神经网络模型。在这里,我们测试的能力,这两种方法来预测尖峰活动在V1清醒的猴子的自然图像。我们发现,迁移学习方法的表现与数据驱动方法相似,并且都优于基于V1现有理论的经典线性-非线性和基于小波的特征表示。值得注意的是,使用预先训练的特征空间的迁移学习需要更少的实验时间来实现相同的性能。总之,多层卷积神经网络(CNN)为预测灵长类动物V1对自然图像的神经反应奠定了新的技术水平,并且为对象识别而学习的深度特征比之前所有的滤波器组理论更好地解释了V1计算。这一发现加强了V1模型的必要性,即远离图像域的多个非线性,它支持解释早期视觉皮层的基础上高层次的功能目标的想法。作者摘要预测的感觉神经元的反应,以任意的自然刺激是非常重要的了解他们的功能。可以说,研究最多的皮层区域是初级视觉皮层(V1),已经开发了许多模型来解释其功能。然而,建立在神经生理学家直觉上的最成功的模型仍然无法解释对自然图像的尖峰反应。在这里,我们使用深度卷积神经网络(CNN)对猴子初级视觉皮层(V1)中的尖峰活动进行建模,这种网络在计算机视觉中已经取得了成功。我们都直接训练CNN来拟合数据,并使用经过训练的CNN来解决高级任务(对象分类)。通过这些方法,我们能够超越以前的模型,并提高预测早期视觉神经元对自然图像的反应的最新水平。我们的研究结果有两个重要的意义。首先,由于V1是几个非线性阶段的结果,因此应该对其进行建模。第二,整个视觉通路的功能模型,其中V1是一个早期阶段,不仅占更高的领域,这样的途径,但也提供了有用的表示V1的预测。
Despite great efforts over several decades, our best models of primary visual cortex (V1) still predict spiking activity quite poorly when probed with natural stimuli, highlighting our limited understanding of the nonlinear computations in V1. Recently, two approaches based on deep learning have emerged for modeling these nonlinear computations: transfer learning from artificial neural networks trained on object recognition and data-driven convolutional neural network models trained end-to-end on large populations of neurons. Here, we test the ability of both approaches to predict spiking activity in response to natural images in V1 of awake monkeys. We found that the transfer learning approach performed similarly well to the data-driven approach and both outperformed classical linear-nonlinear and wavelet-based feature representations that build on existing theories of V1. Notably, transfer learning using a pre-trained feature space required substantially less experimental time to achieve the same performance. In conclusion, multi-layer convolutional neural networks (CNNs) set the new state of the art for predicting neural responses to natural images in primate V1 and deep features learned for object recognition are better explanations for V1 computation than all previous filter bank theories. This finding strengthens the necessity of V1 models that are multiple nonlinearities away from the image domain and it supports the idea of explaining early visual cortex based on high-level functional goals.Author summary Predicting the responses of sensory neurons to arbitrary natural stimuli is of major importance for understanding their function. Arguably the most studied cortical area is primary visual cortex (V1), where many models have been developed to explain its function. However, the most successful models built on neurophysiologists' intuitions still fail to account for spiking responses to natural images. Here, we model spiking activity in primary visual cortex (V1) of monkeys using deep convolutional neural networks (CNNs), which have been successful in computer vision. We both trained CNNs directly to fit the data, and used CNNs trained to solve a high-level task (object categorization). With these approaches, we are able to outperform previous models and improve the state of the art in predicting the responses of early visual neurons to natural images. Our results have two important implications. First, since V1 is the result of several nonlinear stages, it should be modeled as such. Second, functional models of entire visual pathways, of which V1 is an early stage, do not only account for higher areas of such pathways, but also provide useful representations for V1 predictions.