Pragmatically Informative Image Captioning with Character-Level Inference

Pragmatically Informative Image Captioning with Character-Level Inference
复制标题

DOI:
10.18653/v1/n18-2070
复制
发表时间:
2018-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Reuben Cohn-Gordon;Noah D. Goodman;Christopher Potts
Reuben Cohn-Gordon;Noah D. Goodman;Christopher Potts
中科院分区:
其他
文献类型:
--
作者:
Reuben Cohn-Gordon;Noah D. Goodman;Christopher Potts

文献摘要

被引文献

相似文献

我们将神经图像字幕生成器与理性言语行为 (RSA) 模型相结合,打造出一个实用信息丰富的系统:其目标是生成不仅真实的字幕,而且能够将其输入与类似图像区分开来。之前将 RSA 与神经图像字幕相结合的尝试需要对整个可能的话语集进行标准化的推理。这带来了严重的效率问题,以前是通过对可能话语的一小部分进行采样来解决的。相反,我们通过实现一个 RSA 版本来解决这个问题,该版本在字幕展开期间在字符(“a”、“b”、“c”等)级别进行操作。我们发现,仅通过字符级决策即可获得参考字幕的话语级效果。最后,我们介绍了一种用于测试语用说话者模型性能的自动方法,并表明我们的模型优于非语用基线以及单词级 RSA 字幕器。
We combine a neural image captioner with a Rational Speech Acts (RSA) model to make a system that is pragmatically informative: its objective is to produce captions that are not merely true but also distinguish their inputs from similar images. Previous attempts to combine RSA with neural image captioning require an inference which normalizes over the entire set of possible utterances. This poses a serious problem of efficiency, previously solved by sampling a small subset of possible utterances. We instead solve this problem by implementing a version of RSA which operates at the level of characters (“a”, “b”, “c”, ...) during the unrolling of the caption. We find that the utterance-level effect of referential captions can be obtained with only character-level decisions. Finally, we introduce an automatic method for testing the performance of pragmatic speaker models, and show that our model outperforms a non-pragmatic baseline as well as a word-level RSA captioner.