Self-Distillation for Few-Shot Image Captioning

Self-Distillation for Few-Shot Image Captioning
复制标题

DOI:
10.1109/wacv48630.2021.00059
复制
发表时间:
2021-01
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Xianyu Chen;Ming Jiang;Qi Zhao
Xianyu Chen;Ming Jiang;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Xianyu Chen;Ming Jiang;Qi Zhao

文献摘要

相似文献

大规模图像字幕数据集的开发是昂贵的,而大量的未配对图像和文本语料库可能有助于减少人工注释的工作。在本文中,我们研究了少镜头图像字幕问题,只需要少量的注释图像字幕对。我们提出了一种基于集成的自蒸馏方法,该方法允许使用未配对的图像和字幕来训练图像字幕模型。集成由在每次迭代中使用不同数据样本训练的多个基础模型组成。为了从不成对的图像中学习,我们使用集成生成多个伪标题,并根据它们的置信水平分配不同的权重。为了从不成对的字幕中学习,我们提出了一种简单而有效的基于梯度下降的伪特征生成方法。来自集合的伪标题和伪特征用于在未来的迭代中训练基础模型。所提出的方法是通用的不同的图像字幕模型和数据集。我们的实验证明了显着的性能改进和有意义的字幕生成只有1%的配对训练数据。源代码可在https://github.com/chenxy99/SD-FSIC上获得。
The development of large-scale image-captioning datasets is expensive, while the abundance of unpaired images and text corpus can potentially help reduce the efforts of manual annotation. In this paper, we study the few-shot image captioning problem that only requires a small amount of annotated image-caption pairs. We propose an ensemble- based self-distillation method that allows image captioning models to be trained with unpaired images and captions. The ensemble consists of multiple base models trained with different data samples in each iteration. For learning from unpaired images, we generate multiple pseudo captions with the ensemble and allocate different weights according to their confidence levels. For learning from unpaired captions, we propose a simple yet effective pseudo feature generation method based on Gradient Descent. The pseudo captions and pseudo features from the ensemble are used to train the base models in future iterations. The proposed method is general over different image captioning models and datasets. Our experiments demonstrate significant performance improvements and meaningful captions generated with only 1% of paired training data. Source code is available at https://github.com/chenxy99/SD-FSIC.