From image to language and back again

From image to language and back again
复制标题

DOI:
10.1017/s1351324918000086
复制
发表时间:
2018-04
影响因子:
2.5
通讯作者:
Anya Belz;Tamara L. Berg;Licheng Yu
Anya Belz;Tamara L. Berg;Licheng Yu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Anya Belz;Tamara L. Berg;Licheng Yu

文献摘要

相似文献

在过去十年里,计算机视觉和涉及图像和文本的自然语言处理方面的工作经历了爆炸性的增长,尤其是神经网络革命的推动。本卷汇集了来自该领域不同角落的五篇研究文章:多语言多模式图像描述(Frank等人)、多模式机器翻译(Madhythera等人、Frank等人)、图像字幕生成(Madhythera等人、Tanti等人)、视觉场景理解(Silberer等人)以及高级属性的多模式学习(Sorodc等人)。在这篇文章中,我们涉及所有这些主题,因为我们在三个主要标题下回顾了涉及图像和文本的工作:图像描述(第2节)、基于视觉的指代表达生成(REG)和理解(第3节)以及视觉问答(VQA)(第4节)。
Work in computer vision and natural language processing involving images and text has been experiencing explosive growth over the past decade, with a particular boost coming from the neural network revolution. The present volume brings together five research articles from several different corners of the area: multilingual multimodal image description (Frank et al.), multimodal machine translation (Madhyastha et al., Frank et al.), image caption generation (Madhyastha et al., Tanti et al.), visual scene understanding (Silberer et al.), and multimodal learning of high-level attributes (Sorodoc et al.). In this article, we touch upon all of these topics as we review work involving images and text under the three main headings of image description (Section 2), visually grounded referring expression generation (REG) and comprehension (Section 3), and visual question answering (VQA) (Section 4).