Who’s Waldo? Linking People Across Text and Images

Who’s Waldo? Linking People Across Text and Images
复制标题

DOI:
10.1109/iccv48922.2021.00141
复制
发表时间:
2021-08
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Claire Yuqing Cui;Apoorv Khandelwal;Yoav Artzi;Noah Snavely;Hadar Averbuch-Elor
Claire Yuqing Cui;Apoorv Khandelwal;Yoav Artzi;Noah Snavely;Hadar Averbuch-Elor
中科院分区:
其他
文献类型:
--
作者:
Claire Yuqing Cui;Apoorv Khandelwal;Yoav Artzi;Noah Snavely;Hadar Averbuch-Elor

文献摘要

相似文献

我们提出了一个以人为中心的视觉基础的任务和基准数据集,标题中命名的人和图像中描绘的人之间的联系问题。与之前主要基于对象的视觉基础工作相比,我们的新任务掩盖了标题中的人名,以鼓励在这种图像-标题对上训练的方法专注于上下文线索,例如多个人之间的丰富互动,而不是学习名字和外观之间的关联。为了方便这项任务,我们引入了一个新的数据集,谁的沃尔多,自动挖掘维基共享资源上的图像标题数据。我们提出了一种基于transformer的方法,在这项任务上优于几个强大的基线,并将我们的数据发布给研究社区,以促进考虑视觉和语言的上下文模型的工作。代码和数据可在https://whoswaldo.github.io上获得
We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly object-based, our new task masks out the names of people in captions in order to encourage methods trained on such image–caption pairs to focus on contextual cues, such as the rich interactions between multiple people, rather than learning associations between names and appearances. To facilitate this task, we introduce a new dataset, Who’s Waldo, mined automatically from image–caption data on Wikimedia Commons. We propose a Transformer-based method that outperforms several strong baselines on this task, and release our data to the research community to spur work on contextual models that consider both vision and language. Code and data are available at: https://whoswaldo.github.io