Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D

Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D
复制标题

DOI:
--
复制
发表时间:
2020-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Ankit Goyal;Kaiyu Yang;Dawei Yang;Jia Deng
Ankit Goyal;Kaiyu Yang;Dawei Yang;Jia Deng
中科院分区:
其他
文献类型:
--
作者:
Ankit Goyal;Kaiyu Yang;Dawei Yang;Jia Deng

文献摘要

被引文献

相似文献

理解空间关系(例如,“桌上的膝上型计算机”)对于人类和机器人都是重要的。现有的数据集是不够的,因为它们缺乏大规模,高质量的3D地面实况信息,这对于学习空间关系至关重要。在本文中,我们通过构建Rel 3D来填补这一空白:第一个大规模的,人工注释的数据集,用于在3D中建立空间关系。Rel 3D能够量化3D信息在预测大规模人类数据的空间关系方面的有效性。此外,我们提出了最低限度的对比数据收集-一种新的众包方法,以减少数据集的偏见。我们数据集中的3D场景以最小对比度对出现:一对中的两个场景几乎相同,但空间关系在一个场景中保持,而在另一个场景中失败。我们通过经验验证,最低对比度的例子可以诊断当前关系检测模型的问题,并导致样本有效的训练。代码和数据可在https://github.com/princeton-vl/Rel3D上获得。
Understanding spatial relations (e.g., "laptop on table") in visual input is important for both humans and robots. Existing datasets are insufficient as they lack large-scale, high-quality 3D ground truth information, which is critical for learning spatial relations. In this paper, we fill this gap by constructing Rel3D: the first large-scale, human-annotated dataset for grounding spatial relations in 3D. Rel3D enables quantifying the effectiveness of 3D information in predicting spatial relations on large-scale human data. Moreover, we propose minimally contrastive data collection -- a novel crowdsourcing method for reducing dataset bias. The 3D scenes in our dataset come in minimally contrastive pairs: two scenes in a pair are almost identical, but a spatial relation holds in one and fails in the other. We empirically validate that minimally contrastive examples can diagnose issues with current relation detection models as well as lead to sample-efficient training. Code and data are available at https://github.com/princeton-vl/Rel3D.