K-VQG: Knowledge-aware Visual Question Generation for Common-sense Acquisition

K-VQG: Knowledge-aware Visual Question Generation for Common-sense Acquisition
复制标题

DOI:
10.1109/wacv56688.2023.00438
复制
发表时间:
2022-03
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Kohei Uehara;Tatsuya Harada
Kohei Uehara;Tatsuya Harada
中科院分区:
其他
文献类型:
--
作者:
Kohei Uehara;Tatsuya Harada

文献摘要

相似文献

可视化问题生成(VQG)是一个从图像中生成问题的任务。当人们对一幅图像提出问题时,他们的目标往往是获得一些新知识。然而,现有的VQG的研究主要是从答案或问题类别的问题生成,忽略了知识获取的目标。为了将知识获取的观点引入VQG,我们构建了一个新的知识感知VQG数据集K-VQG。这是第一个大型的、人工注释的数据集,其中关于图像的问题与结构化知识相关。我们还开发了一个新的VQG模型,可以编码和使用知识作为问题的目标。在K-VQG数据集上的实验结果表明,该模型的性能优于现有模型。我们的数据集可在https://uehara-mech.github.io/kvqg上公开获取。
Visual Question Generation (VQG) is a task to generate questions from images. When humans ask questions about an image, their goal is often to acquire some new knowledge. However, existing studies on VQG have mainly addressed question generation from answers or question categories, overlooking the objectives of knowledge acquisition. To introduce a knowledge acquisition perspective into VQG, we constructed a novel knowledge-aware VQG dataset called K-VQG. This is the first large, humanly annotated dataset in which questions regarding images are tied to structured knowledge. We also developed a new VQG model that can encode and use knowledge as the target for a question. The experiment results show that our model outperforms existing models on the K-VQG dataset. Our dataset is publicly available at https://uehara-mech.github.io/kvqg.