Extracting Low-Dimensional Psychological Representations from Convolutional Neural Networks

Extracting Low-Dimensional Psychological Representations from Convolutional Neural Networks
复制标题

DOI:
10.1111/cogs.13226
复制
发表时间:
2023-01-01
期刊:
影响因子:
2.5
通讯作者:
Griffiths, Thomas L.
Griffiths, Thomas L.
中科院分区:
心理学3区
文献类型:
--
作者:
Jha, Aditi;Peterson, Joshua C.;Griffiths, Thomas L.

文献摘要

被引文献

相似文献

卷积神经网络(CNN)在心理学和神经科学中被越来越广泛地用于预测人类大脑对视觉图像的反应。通常,CNN使用通过对图像数据集的广泛训练而学习的数千个特征来表示这些图像。这就提出了一个问题:这些特征中有多少真正需要用来模拟人类行为?在这里,我们试图通过两种方式来估计CNN表征中捕捉人类心理表征所需的维度数量:(1)直接使用人类的相似性判断;(2)间接地,在分类的背景下。在这两种情况下,我们发现CNN表示的低维投影足以预测人类的行为。我们表明,这些低维表示可以很容易地解释,为人们如何表示视觉信息提供了进一步的洞察。一系列对照研究表明,这些发现不是由于我们使用的数据集的大小,而可能是由于CNN表示中出现的特征中的高度冗余。
Convolutional neural networks (CNNs) are increasingly widely used in psychology and neuroscience to predict how human minds and brains respond to visual images. Typically, CNNs represent these images using thousands of features that are learned through extensive training on image datasets. This raises a question: How many of these features are really needed to model human behavior? Here, we attempt to estimate the number of dimensions in CNN representations that are required to capture human psychological representations in two ways: (1) directly, using human similarity judgments and (2) indirectly, in the context of categorization. In both cases, we find that low-dimensional projections of CNN representations are sufficient to predict human behavior. We show that these low-dimensional representations can be easily interpreted, providing further insight into how people represent visual information. A series of control studies indicate that these findings are not due to the size of the dataset we used and may be due to a high level of redundancy in the features appearing in CNN representations.