When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes.

When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes.
复制标题

当猪飞行时:合成和自然场景中的上下文推理。

DOI:
10.1109/iccv48922.2021.00032
复制
发表时间:
2021-10
期刊:
... IEEE International Conference on Computer Vision workshops. IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
Kreiman, Gabriel
Kreiman, Gabriel
中科院分区:
其他
文献类型:
--
作者:
Bomatter, Philipp;Zhang, Mengmi;Karev, Dimitar;Madan, Spandan;Tseng, Claire;Kreiman, Gabriel

文献摘要

参考文献

被引文献

相似文献

上下文对人类和机器视觉都至关重要;例如,空中的物体更有可能是飞机而不是猪。上下文的丰富概念包含了几个方面,包括物理规则、统计共现和相对对象大小等。虽然以前的工作主要集中在从网络上获取的脱离上下文的照片来研究场景上下文,但控制上下文违反的性质和程度一直是一项艰巨的任务。在这里,我们引入了一个多样化的、合成的上下文外数据集(OCD),它对场景上下文进行了细粒度控制。通过利用3D模拟引擎,我们系统地控制了虚拟家庭环境中36个对象类别的重力、对象共现和相对大小。我们进行了一系列实验,以深入了解上下文线索对使用强迫症的人类和机器视觉的影响。我们进行了心理物理学实验,以建立人类对非上下文识别的基准,然后将其与最先进的计算机视觉模型进行比较,以量化两者之间的差距。我们提出了一种情境感知识别转换模型,通过多头注意融合对象和情境信息。我们的模型捕获了上下文推理的有用信息,与强迫症和其他上下文外数据集的基线模型相比,在上下文外条件下实现了人类水平的性能和更好的鲁棒性。所有源代码和数据都可以在https://github.com/kreimanlab/WhenPigsFlyContext上公开获得
Context is of fundamental importance to both human and machine vision; e.g., an object in the air is more likely to be an airplane than a pig. The rich notion of context incorporates several aspects including physics rules, statistical co-occurrences, and relative object sizes, among others. While previous work has focused on crowd-sourced out-of-context photographs from the web to study scene context, controlling the nature and extent of contextual violations has been a daunting task. Here we introduce a diverse, synthetic Out-of-Context Dataset (OCD) with fine-grained control over scene context. By leveraging a 3D simulation engine, we systematically control the gravity, object co-occurrences and relative sizes across 36 object categories in a virtual household environment. We conducted a series of experiments to gain insights into the impact of contextual cues on both human and machine vision using OCD. We conducted psychophysics experiments to establish a human benchmark for out-of-context recognition, and then compared it with state-of-the-art computer vision models to quantify the gap between the two. We propose a context-aware recognition transformer model, fusing object and contextual information via multi-head attention. Our model captures useful information for contextual reasoning, enabling human-level performance and better robustness in out-of-context conditions compared to baseline models across OCD and other out-of-context datasets. All source code and data are publicly available at https://github.com/kreimanlab/WhenPigsFlyContext
DOI: 10.1109/tpami.2017.2699184
发表时间: 2018-04-01
影响因子: 23.6
作者:
Chen, Liang-Chieh;Papandreou, George;Yuille, Alan L.
通讯作者: Yuille, Alan L.
DOI: 10.1016/j.patrec.2011.12.004
发表时间: 2012-05-01
影响因子: 5.1
作者:
Choi, Myung Jin;Torralba, Antonio;Willsky, Alan S.
通讯作者: Willsky, Alan S.
DOI: 10.1167/11.9.9
发表时间: 2011-01-01
期刊: JOURNAL OF VISION
影响因子: 1.8
作者:
Mack, Stephen C.;Eckstein, Miguel P.
通讯作者: Eckstein, Miguel P.
DOI: 10.1109/tpami.2016.2577031
发表时间: 2017-06-01
影响因子: 23.6
作者:
Ren, Shaoqing;He, Kaiming;Sun, Jian
通讯作者: Sun, Jian