ScanQA: 3D Question Answering for Spatial Scene Understanding

ScanQA: 3D Question Answering for Spatial Scene Understanding
复制标题

DOI:
10.1109/cvpr52688.2022.01854
复制
发表时间:
2021-12
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Daich Azuma;Taiki Miyanishi;Shuhei Kurita;M. Kawanabe
Daich Azuma;Taiki Miyanishi;Shuhei Kurita;M. Kawanabe
中科院分区:
其他
文献类型:
--
作者:
Daich Azuma;Taiki Miyanishi;Shuhei Kurita;M. Kawanabe

文献摘要

被引文献

相似文献

我们提出了一个新的三维空间理解任务的三维问答(3D-QA)。在3D-QA任务中,模型从丰富的RGB-D室内扫描的整个3D场景接收视觉信息,并回答关于3D场景的给定文本问题。与视觉问题回答的2D问题回答不同,传统的2D-QA模型在对象对齐和方向的空间理解方面存在问题,并且在3D-QA中无法从文本问题中定位对象。我们提出了一个用于3D-QA的基线模型,称为ScanQA 11 https://github.com/ATR-DBI/ScanQA,它从3D对象建议和编码的句子嵌入中学习融合的描述符。这个学习的描述符将语言表达与3D扫描的底层几何特征相关联,并促进3D边界框的回归以确定文本问题中的描述对象。我们收集了人工编辑的问题-答案对,其中自由形式的答案基于每个3D场景中的3D对象。我们新的ScanQA数据集包含来自ScanNet数据集的800个室内场景的41 k多个问答对。据我们所知,ScanQA是第一个在3D环境中执行基于对象的问题回答的大规模努力。
We propose a new 3D spatial understanding task for 3D question answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of a rich RGB-D indoor scan and answer given textual questions about the 3D scene. Unlike the 2D-question answering of visual question answering, the conventional 2D-QA models suffer from problems with spatial understanding of object alignment and directions and fail in object localization from the textual questions in 3D-QA. We propose a baseline model for 3D-QA, called the ScanQA11https://github.com/ATR-DBI/ScanQA, which learns a fused descriptor from 3D object proposals and encoded sentence embeddings. This learned descriptor correlates language expressions with the underlying geometric features of the 3D scan and facilitates the regression of 3D bounding boxes to determine the described objects in textual questions. We collected human-edited question-answer pairs with free-form answers grounded in 3D objects in each 3D scene. Our new ScanQA dataset contains over 41k question-answer pairs from 800 indoor scenes obtained from the ScanNet dataset. To the best of our knowledge, ScanQA is the first large-scale effort to perform object-grounded question answering in 3D environments.