SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition

SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition
复制标题

SpatialSense:空间关系识别的对抗性众包基准

DOI:
--
复制
发表时间:
2019
期刊:
IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
Jia Deng
Jia Deng
中科院分区:
--
文献类型:
--
作者:
Kaiyu Yang;Olga Russakovsky;Jia Deng

文献摘要

参考文献

被引文献

相似文献

理解图像中物体之间的空间关系是一项非常具有挑战性的任务。椅子可能位于人的“后面”,即使它出现在图像中人的左侧(取决于人面向的方向)。如果图像中看起来彼此靠近的两个学生实际上可能并不“相邻”,如果他们之间有第三个学生。我们引入了 SpatialSense,这是一个专门用于空间关系识别的数据集,它捕获了广泛的此类挑战,从而可以对计算机视觉技术进行适当的基准测试。 SpatialSense 是通过对抗性众包构建的,其中人类注释者的任务是找到使用简单线索(例如 2D 空间配置或语言先验)难以预测的空间关系。与现有数据集相比,对抗性众包显着减少了数据集偏差,并在长尾中采样了更有趣的关系。在 SpatialSense 上,最先进的识别模型的表现与简单的基线相当,这表明它们依赖于简单的线索,而不是完全推理这项复杂的任务。 SpatialSense 基准测试为提升计算机视觉系统的空间推理能力提供了一条道路。数据集和代码可在 https://github.com/princeton-vl/SpatialSense 获取。
Understanding the spatial relations between objects in images is a surprisingly challenging task. A chair may be "behind" a person even if it appears to the left of the person in the image (depending on which way the person is facing). Two students that appear close to each other in the image may not in fact be "next to" each other if there is a third student between them. We introduce SpatialSense, a dataset specializing in spatial relation recognition which captures a broad spectrum of such challenges, allowing for proper benchmarking of computer vision techniques. SpatialSense is constructed through adversarial crowdsourcing, in which human annotators are tasked with finding spatial relations that are difficult to predict using simple cues such as 2D spatial configuration or language priors. Adversarial crowdsourcing significantly reduces dataset bias and samples more interesting relations in the long tail compared to existing datasets. On SpatialSense, state-of-the-art recognition models perform comparably to simple baselines, suggesting that they rely on straightforward cues instead of fully reasoning about this complex task. The SpatialSense benchmark provides a path forward to advancing the spatial reasoning capabilities of computer vision systems. The dataset and code are available at https://github.com/princeton-vl/SpatialSense.
DOI: 10.1109/icra.2018.8460538
发表时间: 2017-04
期刊: 2018 IEEE International Conference on Robotics and Automation (ICRA)
影响因子: --
作者:
Zhen Zeng;Zheming Zhou;Zhiqiang Sui;O. C. Jenkins
通讯作者: Zhen Zeng;Zheming Zhou;Zhiqiang Sui;O. C. Jenkins