Question Part Relevance and Editing for Cooperative and Context-Aware VQA (C2VQA)
Question Part Relevance and Editing for Cooperative and Context-Aware VQA (C2VQA)
复制标题
协作和上下文感知 VQA (C2VQA) 的问题部分相关性和编辑
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
M. Nappi
中科院分区:
文献类型:
--
作者:
Andeep S. Toor;H. Wechsler;M. Nappi
Visual Question Answering (VQA), a task that requires the ability to provide an answer to a question given an image, has recently become an important benchmark for computer vision. However, current VQA approaches are unable to adequately handle questions that are "irrelevant", such as asking about a cat for an image that has no cat. To date, only one paper has examined the idea of question relevance in VQA, using a binary classification model to assign a relevancy to the entire question / image pair. Truly robust VQA models, however, must not only identify potentially irrelevant questions, but also discover the source of irrelevance and seek to correct it. We therefore introduce two novel problems, question part relevance and question editing, and approaches for solving each problem. In question part relevance, our models go beyond binary question relevance by assigning a classification probability to the portion of the question that is irrelevant. The best question part relevance classifier is later used in question editing to rank possible corrections to the irrelevant portion of given questions. Two custom datasets are developed for these problems using the Visual Genome dataset as a source. Our best models show promising results in these novel tasks over baseline approaches and models adapted from whole-question relevance classification. This work contributes directly to the development of more context-aware and cooperative VQA models, dubbed C2VQA.