Comparing machines and humans on a visual categorization test

Comparing machines and humans on a visual categorization test
复制标题

DOI:
10.1073/pnas.1109168108
复制
发表时间:
2011-10-25
影响因子:
11.1
通讯作者:
Geman, Donald
Geman, Donald
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Fleuret, Francois;Li, Ting;Geman, Donald

文献摘要

被引文献

相似文献

自动化场景解释受益于机器学习的进步,并且受限任务(例如人脸检测)已经在受限设置下以足够的精度解决。然而,机器在从数字图像中提供丰富的自然场景语义描述方面的性能仍然非常有限,并且远远低于人类。在这里,我们量化了特定环境中的这种“语义差距”:我们比较了人类和机器学习在将图像分配到由组成部分的空间排列确定的两个类别中的一个方面的效率。这些图像不是真实的,但类别定义规则反映了真实的图像的组成结构和语义分析所需的“推理”类型。实验表明,人类受试者从少数例子中掌握了分离原则,而计算机程序的错误率波动很大,即使在接触数千个例子后,仍然远远落后于人类。这些观察结果为计算机视觉的当前趋势提供了支持,例如将机器学习与基于零件的建模相结合。
Automated scene interpretation has benefited from advances in machine learning, and restricted tasks, such as face detection, have been solved with sufficient accuracy for restricted settings. However, the performance of machines in providing rich semantic descriptions of natural scenes from digital images remains highly limited and hugely inferior to that of humans. Here we quantify this "semantic gap" in a particular setting: We compare the efficiency of human and machine learning in assigning an image to one of two categories determined by the spatial arrangement of constituent parts. The images are not real, but the category-defining rules reflect the compositional structure of real images and the type of "reasoning" that appears to be necessary for semantic parsing. Experiments demonstrate that human subjects grasp the separating principles from a handful of examples, whereas the error rates of computer programs fluctuate wildly and remain far behind that of humans even after exposure to thousands of examples. These observations lend support to current trends in computer vision such as integrating machine learning with parts-based modeling.