Evaluating Captioning Models using Markov Logic Networks

Evaluating Captioning Models using Markov Logic Networks
复制标题

DOI:
10.1109/bigdata55660.2022.10020793
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Monika Shah;Somdeb Sarkhel;D. Venugopal
Monika Shah;Somdeb Sarkhel;D. Venugopal
中科院分区:
其他
文献类型:
--
作者:
Monika Shah;Somdeb Sarkhel;D. Venugopal

文献摘要

相似文献

字幕生成等多模态问题从整体上推动了人工智能的发展,因为它们需要集成几个关键领域,如计算机视觉、NLP和知识表示。在本文中,我们开发了一种新的方法来评估字幕模型,验证他们使用马尔可夫逻辑网络(MLN)。具体来说,我们从训练数据中编译MLN,并执行概率推理来估计生成的字幕中的不确定性。为了具体化标题,我们利用自然语言推理(NLI)模型的进步,并将标题转换为MLN的查询。此外,我们使用基于注意力的多实例学习模型将视觉上下文添加到MLN分布中,并基于此增强分布评估标题。我们使用MSCOCO在几个最先进的基准上进行了实验,并表明我们的方法可以像需要人类生成字幕的方法一样有效地评估字幕模型。
Multimodal problems such as caption generation advances AI as a whole since they require integration of several key domains such as computer vision, NLP and knowledge representation. In this paper, we develop a new approach to evaluate captioning models by verifying them using Markov Logic Networks (MLNs). Specifically, we compile an MLN from training data and perform probabilistic inference to estimate uncertainty in a generated caption. To reify the caption, we leverage advances in Natural Language Inference (NLI) models and convert a caption into a query for the MLN. Further, we add visual context into the MLN distribution using an attention-based Multiple Instance Learning model and evaluate a caption based on this augmented distribution. We perform experiments using MSCOCO on several state-of-the-art benchmarks and show that our approach can evaluate captioning models just as effectively as methods that require human-generated captions.