Beyond Narrative Description: Generating Poetry from Images by Multi-Adversarial Training

Beyond Narrative Description: Generating Poetry from Images by Multi-Adversarial Training
复制标题

DOI:
10.1145/3240508.3240587
复制
发表时间:
2018-04
期刊:
Proceedings of the 26th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Bei Liu;Jianlong Fu;Makoto P. Kato;Masatoshi Yoshikawa
Bei Liu;Jianlong Fu;Makoto P. Kato;Masatoshi Yoshikawa
中科院分区:
其他
文献类型:
--
作者:
Bei Liu;Jianlong Fu;Makoto P. Kato;Masatoshi Yoshikawa

文献摘要

相似文献

从图像自动生成自然语言引起了广泛的关注。在本文中,我们进一步研究了将诗歌语言(多行)生成到图像以进行自动诗歌创作。这项任务涉及多重挑战,包括从图像中发现诗意线索(例如,绿色中的希望),以及生成诗歌来满足图像的相关性和语言层面的诗意性。为了解决上述挑战,我们通过策略梯度的多对抗训练将诗歌生成任务制定为两个相关的子任务,通过这可以确保跨模态相关性和诗歌语言风格。为了从图像中提取诗意线索,我们建议学习一种深度耦合的视觉诗意嵌入,其中对象、情感的诗意表示\脚注我们在本研究中将可以表达情感和感受的形容词和动词都视为情感词。和图像中的场景可以共同学习。进一步引入两个判别网络来指导诗歌生成,包括多模态判别器和诗歌风格判别器。为了促进研究,我们发布了两个由人类注释者制作的诗歌数据集,它们具有两个不同的属性:1)第一个人类注释的图像到诗歌对数据集(总共 8,292 美元对),2)迄今为止最大的公共英语诗歌语料库数据集(总共 92,265 美元不同的诗歌)。使用8K图像进行了大量的实验,其中随机选取1.5K图像进行评估。客观和主观评价都显示出与最先进的图像生成诗歌方法相比的优越性能。对超过 500 美元的人类受试者(其中 30 名评估者是诗歌专家)进行的图灵测试证明了我们方法的有效性。
Automatic generation of natural language from images has attracted extensive attention. In this paper, we take one step further to investigate generation of poetic language (with multiple lines) to an image for automatic poetry creation. This task involves multiple challenges, including discovering poetic clues from the image (e.g., hope from green), and generating poems to satisfy both relevance to the image and poeticness in language level. To solve the above challenges, we formulate the task of poem generation into two correlated sub-tasks by multi-adversarial training via policy gradient, through which the cross-modal relevance and poetic language style can be ensured. To extract poetic clues from images, we propose to learn a deep coupled visual-poetic embedding, in which the poetic representation from objects, sentiments \footnoteWe consider both adjectives and verbs that can express emotions and feelings as sentiment words in this research. and scenes in an image can be jointly learned. Two discriminative networks are further introduced to guide the poem generation, including a multi-modal discriminator and a poem-style discriminator. To facilitate the research, we have released two poem datasets by human annotators with two distinct properties: 1) the first human annotated image-to-poem pair dataset (with $8,292$ pairs in total), and 2) to-date the largest public English poem corpus dataset (with $92,265$ different poems in total). Extensive experiments are conducted with 8K images, among which 1.5K image are randomly picked for evaluation. Both objective and subjective evaluations show the superior performances against the state-of-the-art methods for poem generation from images. Turing test carried out with over $500$ human subjects, among which 30 evaluators are poetry experts, demonstrates the effectiveness of our approach.