Cross-domain personalized image captioning

Cross-domain personalized image captioning
复制标题

DOI:
10.1007/s11042-019-7441-7
复制
发表时间:
2019-03
影响因子:
3.6
通讯作者:
Cuirong Long;Xiaoshan Yang;Changsheng Xu
Cuirong Long;Xiaoshan Yang;Changsheng Xu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Cuirong Long;Xiaoshan Yang;Changsheng Xu

文献摘要

相似文献

图像字幕的目的是将图像翻译成完整、自然的句子。它既涉及计算机视觉,也涉及自然语言处理。尽管在深度神经网络的快速发展下,图像字幕取得了良好的效果,但在实际应用中,过分追求字幕模型的评价结果,使得生成的文本描述过于保守。有必要增加文本描述的多样性,并考虑到用户喜欢的词汇和写作风格等先验知识。在本文中,我们研究了个性化的图像字幕,它可以生成句子来描述用户自己的故事和生活感受,并用最喜欢的词语表达。此外,我们还提出了跨域个性化图像字幕(CDPIC)来学习适用于不同社交媒体平台的域不变字幕模型。该方法通过嵌入用户ID作为兴趣向量,可以灵活地对用户兴趣进行建模。据我们所知,我们将用户兴趣建模和简单有效的领域不变约束相结合,提出了第一种跨域的个性化图像字幕方法。在Instagram和Lookbook平台的数据集上验证了该方法的有效性。
Image captioning aims to translate an image to a complete and natural sentence. It involves both computer vision and natural language processing. Though image captioning has achieved good results under the rapid development of deep neural networks, excessively pursuing the evaluation results of the captioning models makes the generated text description too conservative in practical applications. It is necessary to increase the diversity of the text description and account for prior knowledge such as the user’s favorite vocabularies and writing styles. In this paper, we study the personalized image captioning which can generate sentences to describe the user’s own story and feelings of life with the most preferred word expression. Moreover, we propose cross-domain personalized image captioning (CDPIC) to learn domain-invariant captioning models which can be applied on different social media platforms. The proposed method can flexibly model user interest by embedding the user ID as an interest vector. To the best of our knowledge, we propose the first cross-domain personalized image captioning approach by combining the user interest modeling and a simple and effective domain-invariant constraint. The effectiveness of the proposed method is verified on datasets from the Instagram and Lookbook platforms.