RegionSpeak: Quick Comprehensive Spatial Descriptions of Complex Images for Blind Users

RegionSpeak: Quick Comprehensive Spatial Descriptions of Complex Images for Blind Users
复制标题

RegionSpeak:为盲人用户提供复杂图像的快速综合空间描述

DOI:
10.1145/2702123.2702437
复制
发表时间:
2015
期刊:
Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems
影响因子:
--
通讯作者:
Jeffrey P. Bigham
Jeffrey P. Bigham
中科院分区:
--
文献类型:
--
作者:
Yu Zhong;Walter S. Lasecki;Erin L. Brady;Jeffrey P. Bigham

文献摘要

被引文献

相似文献

盲人经常从远程资源中寻求视觉问题的答案,然而,通常采用的单图像,单响应模型并不总是保证用户和源之间有足够的带宽。当问题涉及大量信息或空间布局时尤其如此,例如,这个区域有什么地方可以坐,这个工作台上有什么工具,或者这个机器上的按钮是做什么的?我们的RegionSpeak系统通过为盲人用户提供一种可访问的方式来解决这个问题,以(i)通过图像拼接将多张照片中的联合收割机视觉信息结合起来,(ii)快速收集来自人群的标签,用于并行处理所产生的大视觉区域内包含的所有相关对象,以及(iii)然后交互式地探索被标记的对象的空间布局。区域和描述显示在无障碍触摸屏界面上,允许盲人用户交互式地探索其空间布局。我们证明了亚马逊土耳其机器人的工作人员能够快速准确地识别相关区域,并且要求他们一次只描述一个区域,可以对复杂图像进行更全面的描述。RegionSpeak可用于探索所识别区域的空间布局。它还展示了帮助盲人用户回答困难的空间布局问题的广泛潜力。
Blind people often seek answers to their visual questions from remote sources, however, the commonly adopted single-image, single-response model does not always guarantee enough bandwidth between users and sources. This is especially true when questions concern large sets of information, or spatial layout, e.g., where is there to sit in this area, what tools are on this work bench, or what do the buttons on this machine do? Our RegionSpeak system addresses this problem by providing an accessible way for blind users to (i) combine visual information across multiple photographs via image stitching, em (ii) quickly collect labels from the crowd for all relevant objects contained within the resulting large visual area in parallel, and (iii) then interactively explore the spatial layout of the objects that were labeled. The regions and descriptions are displayed on an accessible touchscreen interface, which allow blind users to interactively explore their spatial layout. We demonstrate that workers from Amazon Mechanical Turk are able to quickly and accurately identify relevant regions, and that asking them to describe only one region at a time results in more comprehensive descriptions of complex images. RegionSpeak can be used to explore the spatial layout of the regions identified. It also demonstrates broad potential for helping blind users to answer difficult spatial layout questions.