Deep supervised, but not unsupervised, models may explain IT cortical representation.

Deep supervised, but not unsupervised, models may explain IT cortical representation.
复制标题

DOI:
10.1371/journal.pcbi.1003915
复制
发表时间:
2014-11
影响因子:
4.3
通讯作者:
Kriegeskorte N
Kriegeskorte N
中科院分区:
生物学2区
文献类型:
--
作者:
Khaligh-Razavi SM;Kriegeskorte N

文献摘要

参考文献

被引文献

相似文献

人和非人类灵长类动物中的颞下皮层可供视觉对象识别。计算对象视觉模型虽然不断改进,但尚未达到人类绩效。目前尚不清楚计算模型的内部表示在多大程度上可以解释IT表示形式。在这里,我们研究了广泛的计算模型表示(总共37个),测试其分类性能以及其考虑IT代表几何形状的能力。这些模型包括众所周知的神经科学对象识别模型(例如HMAX,Visnet)以及计算机视觉(例如SIFT,GIST,自相似性特征和深度卷积神经网络)的几种模型。我们比较了模型表示的代表性差异矩阵(RDM)与从人IT(用fMRI测量)获得的RDM,并在相同的刺激中(用细胞记录测量)(未在训练模型中使用)。更好的性能模型与它更相似,因为它们显示出更大的代表模式的聚类。此外,更好的性能模型也更像它们的类别内代表性差异。代表几何形状与许多模型之间显着相关。但是,在其中观察到的分类聚类在很大程度上无法解释。深度卷积网络通过监督培训,并获得了超过一百万个标签的图像,但达到了最高的分类性能,并且最好地解释了它,尽管它没有完全解释IT数据。将该模型的特征与适当的权重结合,并添加线性组合,以最大程度地提高动画和无生命的对象之间的边距以及面部和其他对象之间的余量,产生了一个充分解释我们IT数据的表示。总体而言,我们的结果表明,解释它需要通过监督学习训练的计算特征,以强调其在其中突出反映在行为上重要的分类部门。 计算机还不能像人类一样识别对象。计算机视觉可能会从生物愿景中学习。但是,神经科学尚未解释大脑如何识别对象,并且必须从计算机视觉中汲取初始计算模型。为了解决这个鸡和蛋的问题,我们将37个计算模型表示与生物学大脑的表示。模型表示越相似于高级视觉脑表示,模型在对象分类下执行的越好。大多数模型都没有解释大脑表示,因为它们错过了动画和无动物之间以及面部和其他物体之间在灵长类动物大脑中突出的对象。一个深层神经网络模型,通过监督培训,并具有超过100万个标记的图像,并代表计算机视觉中的最新状态,最接近解释大脑表示。我们的大脑似乎将视觉输入强加于某些对成功行为很重要的分类分区。大脑可能会通过进化和个人经验来学习这些分裂。计算机视觉类似地需要使用许多标记的图像进行学习,以强调正确的分类分区。
Inferior temporal (IT) cortex in human and nonhuman primates serves visual object recognition. Computational object-vision models, although continually improving, do not yet reach human performance. It is unclear to what extent the internal representations of computational models can explain the IT representation. Here we investigate a wide range of computational model representations (37 in total), testing their categorization performance and their ability to account for the IT representational geometry. The models include well-known neuroscientific object-recognition models (e.g. HMAX, VisNet) along with several models from computer vision (e.g. SIFT, GIST, self-similarity features, and a deep convolutional neural network). We compared the representational dissimilarity matrices (RDMs) of the model representations with the RDMs obtained from human IT (measured with fMRI) and monkey IT (measured with cell recording) for the same set of stimuli (not used in training the models). Better performing models were more similar to IT in that they showed greater clustering of representational patterns by category. In addition, better performing models also more strongly resembled IT in terms of their within-category representational dissimilarities. Representational geometries were significantly correlated between IT and many of the models. However, the categorical clustering observed in IT was largely unexplained by the unsupervised models. The deep convolutional network, which was trained by supervision with over a million category-labeled images, reached the highest categorization performance and also best explained IT, although it did not fully explain the IT data. Combining the features of this model with appropriate weights and adding linear combinations that maximize the margin between animate and inanimate objects and between faces and other objects yielded a representation that fully explained our IT data. Overall, our results suggest that explaining IT requires computational features trained through supervised learning to emphasize the behaviorally important categorical divisions prominently reflected in IT. Computers cannot yet recognize objects as well as humans can. Computer vision might learn from biological vision. However, neuroscience has yet to explain how brains recognize objects and must draw from computer vision for initial computational models. To make progress with this chicken-and-egg problem, we compared 37 computational model representations to representations in biological brains. The more similar a model representation was to the high-level visual brain representation, the better the model performed at object categorization. Most models did not come close to explaining the brain representation, because they missed categorical distinctions between animates and inanimates and between faces and other objects, which are prominent in primate brains. A deep neural network model that was trained by supervision with over a million category-labeled images and represents the state of the art in computer vision came closest to explaining the brain representation. Our brains appear to impose upon the visual input certain categorical divisions that are important for successful behavior. Brains might learn these divisions through evolution and individual experience. Computer vision similarly requires learning with many labeled images so as to emphasize the right categorical divisions.
DOI: 10.1561/2200000006
发表时间: 2009-01-01
影响因子: 32.8
作者:
Bengio, Yoshua
通讯作者: Bengio, Yoshua
DOI: 10.1523/jneurosci.3809-13.2013
发表时间: 2013-11-27
期刊: The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子: --
作者:
Devereux BJ;Clarke A;Marouchos A;Tyler LK
通讯作者: Tyler LK
DOI: 10.1152/jn.90657.2008
发表时间: 2009-02-01
影响因子: 2.5
作者:
Bell, Andrew H.;Hadj-Bouziane, Fadila;Ungerleider, Leslie G.
通讯作者: Ungerleider, Leslie G.
DOI: 10.1523/jneurosci.2828-13.2014
发表时间: 2014-04-02
影响因子: 5.3
作者:
Clarke, Alex;Tyler, Lorraine K.
通讯作者: Tyler, Lorraine K.
DOI: 10.1371/journal.pcbi.1003167
发表时间: 2013
影响因子: 4.3
作者:
Baldassi C;Alemi-Neissi A;Pagan M;Dicarlo JJ;Zecchina R;Zoccolan D
通讯作者: Zoccolan D