VizWiz Grand Challenge: Answering Visual Questions from Blind People

VizWiz Grand Challenge: Answering Visual Questions from Blind People
复制标题

DOI:
10.1109/cvpr.2018.00380
复制
发表时间:
2018-02
期刊:
2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham
D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham
中科院分区:
其他
文献类型:
--
作者:
D. Gurari;Qing Li;Abigale Stangl;Anhong Guo;Chi Lin;K. Grauman;Jiebo Luo;Jeffrey P. Bigham

文献摘要

被引文献

相似文献

算法以自动回答视觉问题的算法是由人工VQA设置中构建的视觉问题回答(VQA)数据集的动机。我们提出了Vizwiz,这是第一个面向目标的VQA数据集,该数据集是由自然VQA设置产生的。 Vizwiz由31,000多个视觉问题组成,这些问题来自盲人,他们每个人都使用手机拍照,并记录了有关此问题的口语问题,每个视觉问题也有10个众包答案。 Vizwiz与许多现有的VQA数据集有所不同,因为(1)图像是由盲目摄影师捕获的,因此质量差,(2)问题是在说话,因此更加对话,并且(3)通常无法回答视觉问题。评估用于回答视觉问题并确定视觉问题是否可以回答的现代算法的评估表明,Vizwiz是一个具有挑战性的数据集。我们介绍此数据集,以鼓励更大的社区开发更普遍的算法,以帮助盲人。
The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people.