Learning Convolutional Text Representations for Visual Question Answering

Learning Convolutional Text Representations for Visual Question Answering
复制标题

DOI:
10.1137/1.9781611975321.67
复制
发表时间:
2017-05
期刊:
--
影响因子:
--
通讯作者:
Zhengyang Wang;Shuiwang Ji
Zhengyang Wang;Shuiwang Ji
中科院分区:
其他
文献类型:
--
作者:
Zhengyang Wang;Shuiwang Ji

文献摘要

被引文献

相似文献

视觉问答是最近提出的一项人工智能任务,需要对图像和文本都有深刻的理解。在深度学习中,图像通常通过卷积神经网络建模,文本通常通过循环神经网络建模。虽然对图像建模的要求与传统的计算机视觉任务(如对象识别和图像分类)相似,但与其他自然语言处理任务相比,视觉问答对文本表示提出了不同的需求。在这项工作中,我们对视觉问答中的自然语言问题进行了详细的分析。在此基础上,我们提出依靠卷积神经网络来学习文本表示。通过探索专门用于文本数据的卷积神经网络的各种属性,如宽度和深度,我们提出了我们的“CNN Inception + Gate”模型。我们表明,我们的模型提高了问题表示,从而提高了视觉问答模型的整体准确性。我们还发现,视觉问答对文本表示的要求比传统的自然语言处理任务更复杂和全面,这使得它成为评估文本表示方法的更好任务。像fastText这样的浅层模型在文本分类等任务中可以获得与深度学习模型相当的结果,但不适合用于视觉问答。
Visual question answering is a recently proposed artificial intelligence task that requires a deep understanding of both images and texts. In deep learning, images are typically modeled through convolutional neural networks, and texts are typically modeled through recurrent neural networks. While the requirement for modeling images is similar to traditional computer vision tasks, such as object recognition and image classification, visual question answering raises a different need for textual representation as compared to other natural language processing tasks. In this work, we perform a detailed analysis on natural language questions in visual question answering. Based on the analysis, we propose to rely on convolutional neural networks for learning textual representations. By exploring the various properties of convolutional neural networks specialized for text data, such as width and depth, we present our "CNN Inception + Gate" model. We show that our model improves question representations and thus the overall accuracy of visual question answering models. We also show that the text representation requirement in visual question answering is more complicated and comprehensive than that in conventional natural language processing tasks, making it a better task to evaluate textual representation methods. Shallow models like fastText, which can obtain comparable results with deep learning models in tasks like text classification, are not suitable in visual question answering.