Knowledge Acquisition for Human-In-The-Loop Image Captioning

Knowledge Acquisition for Human-In-The-Loop Image Captioning
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Ervine Zheng;Qi Yu;Rui Li;Pengcheng Shi;Anne R. Haake
Ervine Zheng;Qi Yu;Rui Li;Pengcheng Shi;Anne R. Haake
中科院分区:
其他
文献类型:
--
作者:
Ervine Zheng;Qi Yu;Rui Li;Pengcheng Shi;Anne R. Haake

文献摘要

相似文献

图像字幕提供了一个计算过程来理解图像的语义,并使用描述性语言来传达它们。然而,由于图像的复杂性质和训练数据的质量/大小,自动字幕模型可能不总是生成令人满意的字幕。我们提出了一个交互式字幕框架,通过让人类处于循环中并执行线上-线下知识获取(KA)过程来改进机器生成的字幕。具体地说,在线KA接受人类用户指定的关键字列表,并将它们与图像特征融合以生成捕捉图像语义的可读句子。它利用多模式条件字幕完成机制来确保所有用户输入的关键字都出现在生成的字幕中。离线KA进一步从用户输入中学习以更新模型,并在将来为看不见的图像生成字幕。它建立在贝叶斯转换器架构的基础上,该架构动态分配神经资源,并支持不确定性感知模型更新,以缓解过度匹配。我们的理论分析也证明了离线KA会自动选择最优的模型容量来容纳新获得的知识。在真实世界数据上的实验证明了该框架的有效性。
Image captioning offers a computational process to understand the semantics of images and convey them using descriptive language. However, automated captioning models may not always generate satisfactory captions due to the complex nature of the images and the quality/size of the training data. We propose an interactive captioning framework to improve machine-generated cap-tions by keeping humans in the loop and performing an online-offline knowledge acquisition (KA) process. In particular, online KA accepts a list of keywords specified by human users and fuses them with the image features to generate a read-able sentence that captures the semantics of the image. It leverages a multimodal conditioned caption completion mechanism to ensure the appearance of all user-input keywords in the generated caption. Offline KA further learns from the user inputs to update the model and benefits caption generation for unseen images in the future. It is built upon a Bayesian transformer architecture that dynamically allocates neural resources and supports uncertainty-aware model updates to mitigate overfitting. Our theoretical analysis also proves that Offline KA automatically selects the best model capacity to accommodate the newly acquired knowledge. Experiments on real-world data demonstrate the effectiveness of the proposed framework.