Visual scenes are categorized by function.

Visual scenes are categorized by function.
复制标题

DOI:
10.1037/xge0000129
复制
发表时间:
2016-01
期刊:
Journal of experimental psychology. General
影响因子:
--
通讯作者:
Fei-Fei L
Fei-Fei L
中科院分区:
其他
文献类型:
--
作者:
Greene MR;Baldassano C;Esteva A;Beck DM;Fei-Fei L

文献摘要

被引文献

相似文献

我们怎么知道一个厨房是一个厨房看?传统的模型认为场景分类是通过识别必要和充分的特征和对象来实现的,但对于这些特征和对象的定义却没有一致的认识。然而,场景类别应该反映我们如何使用视觉信息。因此,我们测试的假设,场景类别反映功能,或在一个场景中的动作的可能性。我们的方法是比较人类的分类模式与预测的功能和替代模型。我们收集了一个大规模的场景类别距离矩阵(500万次试验),要求观察者简单地决定两张图像是来自相同还是不同的类别。使用美国时间使用调查中的动作,我们将动作映射到每个场景(140万次试验)。我们发现,排名类别距离和功能距离(r=0.50,或66%的最大可能的相关性)之间有很强的关系。函数模型优于基于对象的距离(r=0.33),卷积神经网络的视觉特征(r=0.39),词汇距离(r=0.27)和视觉特征模型的替代模型。使用分层线性回归,我们发现函数捕获了85.5%的总体解释方差,其中近一半的解释方差仅由函数捕获,这意味着替代模型的预测能力是由于它们与基于函数的模型共享方差。这些结果挑战了主流学派的思想,视觉特征和对象是足够的场景分类,而不是一个场景的类别可能是由场景的功能。
How do we know that a kitchen is a kitchen by looking? Traditional models posit that scene categorization is achieved through recognizing necessary and sufficient features and objects, yet there is little consensus about what these may be. However, scene categories should reflect how we use visual information. We therefore test the hypothesis that scene categories reflect functions, or the possibilities for actions within a scene. Our approach is to compare human categorization patterns with predictions made by both functions and alternative models. We collected a large-scale scene category distance matrix (5 million trials) by asking observers to simply decide whether two images were from the same or different categories. Using the actions from the American Time Use Survey, we mapped actions onto each scene (1.4 million trials). We found a strong relationship between ranked category distance and functional distance (r=0.50, or 66% of the maximum possible correlation). The function model outperformed alternative models of object-based distance (r=0.33), visual features from a convolutional neural network (r=0.39), lexical distance (r=0.27), and models of visual features. Using hierarchical linear regression, we found that functions captured 85.5% of overall explained variance, with nearly half of the explained variance captured only by functions, implying that the predictive power of alternative models was due to their shared variance with the function-based model. These results challenge the dominant school of thought that visual features and objects are sufficient for scene categorization, suggesting instead that a scene’s category may be determined by the scene’s function.