Visual Question Answer Diversity

Visual Question Answer Diversity
复制标题

DOI:
10.1609/hcomp.v6i1.13341
复制
发表时间:
2018-06
期刊:
--
影响因子:
--
通讯作者:
Chun-Ju Yang;K. Grauman;D. Gurari
Chun-Ju Yang;K. Grauman;D. Gurari
中科院分区:
其他
文献类型:
--
作者:
Chun-Ju Yang;K. Grauman;D. Gurari

文献摘要

被引文献

相似文献

视觉问题(VQ)可以导致多个人用不同的答案来回答,而不是单一的、商定的回答。此外,来自人群的答案可以包括不同数量的唯一答案,这些答案以不同的相对频率出现。出现这种答案差异的原因有很多,包括VQ是主观的、困难的或模棱两可的。我们提出了一个新的问题来预测任何给定的VQ从人群中观察到的答案分布,即唯一答案的数量及其相对频率。我们的实验证实,对于盲人和有视力的人提出的VQ,答案分布都可以被准确地预测。然后,我们提出了一种新颖的大众支持的VQA系统,该系统使用答案分布预测来推理需要多少答案才能捕捉到可能的人类反应的多样性。实验表明,与最先进的系统相比,该系统加快了捕获答案多样性的速度,所需的人力资源要少得多。
Visual questions (VQs) can lead multiple people to respond with different answers rather than a single, agreed upon response. Moreover, the answers from a crowd can include different numbers of unique answers that arise with different relative frequencies. Such answer diversity arises for a variety of reasons including that VQs are subjective, difficult, or ambiguous. We propose a new problem of predicting the answer distribution that would be observed from a crowd for any given VQ; i.e., the number of unique answers and their relative frequencies. Our experiments confirm that the answer distribution can be predicted accurately for VQs asked by both blind and sighted people. We then propose a novel crowd-powered VQA system that uses the answer distribution predictions to reason about how many answers are needed to capture the diversity of possible human responses. Experiments demonstrate this proposed system accelerates capturing the diversity of answers with considerably less human effort than is required with a state-of-art system.