Multi-task Learning for Captioning Images with Novel Words

Multi-task Learning for Captioning Images with Novel Words
复制标题

用新词为图像添加字幕的多任务学习

DOI:
10.1049/iet-cvi.2018.5005
复制
发表时间:
--
影响因子:
1.7
通讯作者:
Xuzhi Li
Xuzhi Li
中科院分区:
计算机科学4区
文献类型:
--
作者:
He Zheng;Jiahong Wu;Rui Liang;Ye Li;Xuzhi Li

文献摘要

相似文献

最近的字幕模型在描述成对图像-句子对中看不到的概念的能力方面是有限的。这项研究提出了一个多任务学习的框架,用于描述现有图像字幕数据集中不存在的新单词。作者的框架利用了来自图像分类数据集的外部源标记图像,以及从注释文本中提取的语义知识。他们提出最小化一个联合目标,该目标可以从这些不同的数据源中学习,并利用分布式语义嵌入。在推理步骤中,他们通过考虑字幕模型和语言模型来改变BeamSearch步骤,使模型能够概括图像字幕数据集之外的新单词。他们证明,在框架中添加一个注释的文本数据,可以帮助图像字幕模型描述图像与正确的相应的新颖的话。在两种不同语言的AI Challenger和Microsoft coco(MSCOCO)图像字幕数据集上进行了广泛的实验,证明了其框架描述场景和对象等新颖词语的能力。
Recent captioning models are limited in their ability to describe concepts unseen in paired image–sentence pairs. This study presents a framework of multi‐task learning for describing novel words not present in existing image‐captioning datasets. The authors’ framework takes advantage of external sources‐labelled images from image classification datasets, and semantic knowledge extracted from the annotated text. They propose minimising a joint objective which can learn from these diverse data sources and leverage distributional semantic embeddings. When in the inference step they change the BeamSearch step by considering both the caption model and language model enabling the model to generalise novel words outside of image‐captioning datasets. They demonstrate that in the framework by adding an annotated text data which can help the image captioning model to describe images with the right corresponding novel words. Extensive experiments are conducted on both AI Challenger and Microsoft coco (MSCOCO) image captioning datasets of two different languages, demonstrating the ability of their framework to describe novel words such as scenes and objects.