Leveraging Human Attention in Novel Object Captioning

Leveraging Human Attention in Novel Object Captioning
复制标题

DOI:
10.24963/ijcai.2021/86
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Xianyu Chen;Ming Jiang;Qi Zhao
Xianyu Chen;Ming Jiang;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Xianyu Chen;Ming Jiang;Qi Zhao

文献摘要

相似文献

图像字幕模型依赖于成对的图像-文本语料库的训练,这在描述包含训练数据中不存在的新对象的图像时提出了各种挑战。虽然以前的新对象字幕方法依赖于外部图像标记器或对象检测器来描述新对象,但我们提出了基于注意力的新对象字幕(ANOC),它补充了具有人类注意力特征的新对象字幕,这些特征通常是独立于任务的重要信息。它引入了一种门控机制,自适应地将人类注意力与自学习的机器注意力相结合,并采用约束自批判序列训练方法来解决曝光偏差,同时保持对新对象描述的约束。在nocaps和Held-Out COCO数据集上进行的大量实验表明,我们的方法大大优于最先进的新型对象字幕。我们的源代码可在https://github.com/chenxy99/ANOC上获得。
Image captioning models depend on training with paired image-text corpora, which poses various challenges in describing images containing novel objects absent from the training data. While previous novel object captioning methods rely on external image taggers or object detectors to describe novel objects, we present the Attention-based Novel Object Captioner (ANOC) that complements novel object captioners with human attention features that characterize generally important information independent of tasks. It introduces a gating mechanism that adaptively incorporates human attention with self-learned machine attention, with a Constrained Self-Critical Sequence Training method to address the exposure bias while maintaining constraints of novel object descriptions. Extensive experiments conducted on the nocaps and Held-Out COCO datasets demonstrate that our method considerably outperforms the state-of-the-art novel object captioners. Our source code is available at https://github.com/chenxy99/ANOC.