Question Part Relevance and Editing for Cooperative and Context-Aware VQA (C2VQA)

Question Part Relevance and Editing for Cooperative and Context-Aware VQA (C2VQA)
复制标题

协作和上下文感知 VQA (C2VQA) 的问题部分相关性和编辑

DOI:
--
复制
发表时间:
2017
期刊:
International Conference on Content-Based Multimedia Indexing
影响因子:
--
通讯作者:
M. Nappi
M. Nappi
中科院分区:
--
文献类型:
--
作者:
Andeep S. Toor;H. Wechsler;M. Nappi

文献摘要

被引文献

相似文献

视觉问答(VQA)是一项需要能够对给定图像的问题提供答案的任务,最近已成为计算机视觉的重要基准。然而,当前的VQA方法无法充分处理“不相关”的问题,例如向没有猫的图像询问猫。迄今为止,只有一篇论文研究了VQA中问题相关性的想法,使用二元分类模型来分配整个问题/图像对的相关性。然而,真正强大的VQA模型不仅要识别潜在的不相关问题,还要发现不相关的来源并寻求纠正它。因此,我们介绍了两个新的问题,问题部分的相关性和问题编辑,以及解决每个问题的方法。在问题部分的相关性,我们的模型超越了二元问题的相关性,通过分配一个分类概率的问题是不相关的部分。最佳问题部分相关性分类器稍后在问题编辑中用于对给定问题的不相关部分的可能校正进行排名。使用Visual Genome数据集作为源,为这些问题开发了两个自定义数据集。我们最好的模型在这些新任务中显示出了良好的效果,超过了基线方法和从整个问题相关性分类中改编的模型。这项工作直接有助于开发更多的上下文感知和合作的VQA模型,称为C2 VQA。
Visual Question Answering (VQA), a task that requires the ability to provide an answer to a question given an image, has recently become an important benchmark for computer vision. However, current VQA approaches are unable to adequately handle questions that are "irrelevant", such as asking about a cat for an image that has no cat. To date, only one paper has examined the idea of question relevance in VQA, using a binary classification model to assign a relevancy to the entire question / image pair. Truly robust VQA models, however, must not only identify potentially irrelevant questions, but also discover the source of irrelevance and seek to correct it. We therefore introduce two novel problems, question part relevance and question editing, and approaches for solving each problem. In question part relevance, our models go beyond binary question relevance by assigning a classification probability to the portion of the question that is irrelevant. The best question part relevance classifier is later used in question editing to rank possible corrections to the irrelevant portion of given questions. Two custom datasets are developed for these problems using the Visual Genome dataset as a source. Our best models show promising results in these novel tasks over baseline approaches and models adapted from whole-question relevance classification. This work contributes directly to the development of more context-aware and cooperative VQA models, dubbed C2VQA.