Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision
复制标题

DOI:
10.1109/iccv48922.2021.00048
复制
发表时间:
2021-08
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Xiaoshi Wu;Hadar Averbuch-Elor;J. Sun;Noah Snavely
Xiaoshi Wu;Hadar Averbuch-Elor;J. Sun;Noah Snavely
中科院分区:
其他
文献类型:
--
作者:
Xiaoshi Wu;Hadar Averbuch-Elor;J. Sun;Noah Snavely

文献摘要

相似文献

过去二十年来,地标和城市的互联网照片的丰富性导致了 3D 视觉领域的重大进步,包括根据旅游照片自动 3D 重建世界地标。然而,这些 3D 增强馆藏可用的主要信息来源(即语言,例如来自图像说明的语言)实际上尚未开发。在这项工作中,我们提出了 WikiScenes,这是一个新的大规模地标照片集数据集,其中包含标题和分层类别名称形式的描述性文本。 WikiScenes 形成了涉及图像、文本和 3D 几何的多模态推理的新测试平台。我们展示了 WikiScenes 在学习图像和 3D 模型上的语义概念方面的实用性。我们的弱监督框架连接图像、3D 结构和语义——利用 3D 几何提供的强约束——将语义概念与图像像素和 3D 点关联起来。1
The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world’s landmarks from tourist photos. However, a major source of information available for these 3D-augmented collections—namely language, e.g., from image captions—has been virtually untapped. In this work, we present WikiScenes, a new, large-scale dataset of landmark photo collections that contains descriptive text in the form of captions and hierarchical category names. WikiScenes forms a new testbed for multimodal reasoning involving images, text, and 3D geometry. We demonstrate the utility of WikiScenes for learning semantic concepts over images and 3D models. Our weakly-supervised framework connects images, 3D structure, and semantics—utilizing the strong constraints provided by 3D geometry—to associate semantic concepts to image pixels and 3D points.1