Improving Image Captioning with Conditional Generative Adversarial Nets

Improving Image Captioning with Conditional Generative Adversarial Nets
复制标题

DOI:
10.1609/aaai.v33i01.33018142
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Chen Chen-Chen;Shuai Mu;Wanpeng Xiao;Zexiong Ye;Liesi Wu;Fuming Ma;Qi Ju
Chen Chen-Chen;Shuai Mu;Wanpeng Xiao;Zexiong Ye;Liesi Wu;Fuming Ma;Qi Ju
中科院分区:
其他
文献类型:
--
作者:
Chen Chen-Chen;Shuai Mu;Wanpeng Xiao;Zexiong Ye;Liesi Wu;Fuming Ma;Qi Ju

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的条件生成对抗网络为基础的图像字幕框架作为传统的基于学习(RL)的编码器-解码器架构的扩展。为了解决不同客观语言度量之间的不一致评价问题,我们设计了一些“自适应”网络来自动和渐进地确定生成的字幕是人类描述的还是机器生成的。介绍了两种神经网络结构(CNN和基于RNN的结构),因为每种结构都有自己的优点。该算法是通用的,因此它可以增强任何现有的基于RL的图像字幕框架,我们表明,传统的RL训练方法只是我们的方法的一个特例。从经验上讲,我们在所有语言评估指标上表现出一致的改进,用于不同的最先进的图像字幕模型。此外,经过良好训练的鉴别器也可以被看作是客观的图像字幕评价器。
In this paper, we propose a novel conditional-generativeadversarial-nets-based image captioning framework as an extension of traditional reinforcement-learning (RL)-based encoder-decoder architecture. To deal with the inconsistent evaluation problem among different objective language metrics, we are motivated to design some “discriminator” networks to automatically and progressively determine whether generated caption is human described or machine generated. Two kinds of discriminator architectures (CNN and RNNbased structures) are introduced since each has its own advantages. The proposed algorithm is generic so that it can enhance any existing RL-based image captioning framework and we show that the conventional RL training method is just a special case of our approach. Empirically, we show consistent improvements over all language evaluation metrics for different state-of-the-art image captioning models. In addition, the well-trained discriminators can also be viewed as objective image captioning evaluators.