NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External Knowledge

NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External Knowledge
复制标题

NOC-REK:从外部知识检索词汇的新颖对象描述

DOI:
10.1109/cvpr52688.2022.01747
复制
发表时间:
2022
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Hideki Nakayama
Hideki Nakayama
中科院分区:
--
文献类型:
--
作者:
Duc Minh Vo;Hong Chen;Akihiro Sugimoto;Hideki Nakayama

文献摘要

相似文献

新颖的对象描述旨在描述训练数据中缺少的对象,其关键要素是为模型提供对象词汇。尽管现有方法严重依赖于对象检测模型,但我们将检测步骤视为从外部知识进行词汇检索,其形式为维基词典中任何对象定义的嵌入形式,其中我们在检索图像区域中使用从 Transformer 模型学习的特征。我们提出了一种从外部知识中检索词汇的端到端新颖对象字幕方法(NOC-REK),该方法同时学习词汇检索和字幕生成,成功地描述了训练数据集之外的新颖对象。此外,我们的模型通过在新对象出现时简单地更新外部知识来消除模型重新训练的要求。我们对 COCO 和 Nocaps 数据集的综合实验表明,我们的 NOCREK 对 SOTA 相当有效。
Novel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection model, we view the detection step as vocabulary retrieval from an external knowledge in the form of embeddings for any object's definition from Wiktionary, where we use in the retrieval image region features learned from a transformers model. We propose an end-to-end Novel Object Captioning with Retrieved vocabulary from External Knowledge method (NOC-REK), which simultaneously learns vocabulary retrieval and caption generation, successfully describing novel objects outside of the training dataset. Furthermore, our model eliminates the requirement for model retraining by simply updating the external knowledge whenever a novel object appears. Our comprehensive experiments on held-out COCO and Nocaps datasets show that our NOCREK is considerably effective against SOTAs.