Multiple Perspective Caption Generation with Attention Mechanism

Multiple Perspective Caption Generation with Attention Mechanism
复制标题

DOI:
10.1109/iiai-aai50415.2020.00031
复制
发表时间:
2020-09
期刊:
2020 9th International Congress on Advanced Applied Informatics (IIAI-AAI)
影响因子:
--
通讯作者:
H. Yanagimoto;Maaki Shozu
H. Yanagimoto;Maaki Shozu
中科院分区:
其他
文献类型:
--
作者:
H. Yanagimoto;Maaki Shozu

文献摘要

相似文献

在字幕生成中,字幕生成系统生成字幕,该字幕用自然语言描述图像的内容,并且需要理解图像和文本。因此,字幕生成是自然语言处理和图像处理中的一项重要任务.近年来,由于深度学习可以构造图像处理和自然语言处理所共有的中间表示,因此,许多研究者将其作为构建字幕生成系统的关键技术。首先,系统使用卷积神经网络从给定图像生成特征。最终,系统根据该特征生成单词序列、标题。这意味着该系统由两个模块组成,图像处理模块和语言模型模块,并且这两个模块同时使用训练数据集进行训练。基于深度学习的字幕生成系统是一个黑盒系统,很难与人类合作。因此,我们在字幕生成系统中引入了注意力机制,通过注意力权重来控制字幕的生成。由于通常的字幕生成系统可以直接从图像生成字幕,因此该系统只能从一个图像生成单个字幕。由于我们可以控制初始的注意力权重,因此该系统可以从同一幅图像中生成不同的字幕。我们的研究结果展示了我们如何与深度学习合作,以及合作改善了字幕生成的事实。
In caption generation, a caption generation system generates a caption, which describes the content of the image with natural language and needs to understand both an image and a text. So caption generation is an essential task in natural language processing and image processing. Many researchers recently pay attention to deep learning as a key technique to construct the caption generation system because deep learning can construct an intermediate representation, which is shared in both image processing and natural language processing. First, the system generates a feature from a given image with convolutional neural networks. Eventually, the system generates a word sequence, a caption, from the feature. It means that the system consists of two modules, the image processing module and the language model module and both of the modules are simultaneously trained with a training dataset. The deep learning based caption generation system is a blackbox system and it is difficult to collaborate with a human. So, we introduce attention mechanism in the caption generation system and we control caption generation by the attention weights. The usual caption generation systems can generate only a single caption from one image because the system can generate a caption from an image directly. The proposed system can generate some different captions form the same image because we can control the initial attention weights. Our results demonstrate how we collaborate with deep learning and the fact that the collaboration improves caption generation.