VizWiz-FewShot: Locating Objects in Images Taken by People With Visual Impairments

VizWiz-FewShot: Locating Objects in Images Taken by People With Visual Impairments
复制标题

DOI:
10.48550/arxiv.2207.11810
复制
发表时间:
2022-07
期刊:
--
影响因子:
--
通讯作者:
Yu-Yun Tseng;Alexander Bell;D. Gurari
Yu-Yun Tseng;Alexander Bell;D. Gurari
中科院分区:
其他
文献类型:
--
作者:
Yu-Yun Tseng;Alexander Bell;D. Gurari

文献摘要

相似文献

我们介绍了一个几张照片的定位数据集,这些数据集来自摄影师,他们真实地试图了解他们拍摄的图像中的视觉内容。它包括视力障碍者拍摄的4500多张图像中的100个类别的近10,000个分段。与现有的少镜头目标检测和实例分割数据集相比,我们的数据集是第一个定位对象中的洞的数据集(例如,在12.3%的分割中发现),它显示了相对于图像占据更大大小范围的对象,并且文本在对象中的常见程度超过五倍(例如,在22.4%的分割中)。对三种现代少镜头定位算法的分析表明,它们对我们的新数据集的泛化能力很差。这些算法通常很难定位有洞的对象、非常小的和非常大的对象以及缺少文本的对象。为了鼓励更大的社区来解决这些悬而未决的挑战,我们在https://vizwiz.org上公开分享了我们的注释少的数据集。
We introduce a few-shot localization dataset originating from photographers who authentically were trying to learn about the visual content in the images they took. It includes nearly 10,000 segmentations of 100 categories in over 4,500 images that were taken by people with visual impairments. Compared to existing few-shot object detection and instance segmentation datasets, our dataset is the first to locate holes in objects (e.g., found in 12.3\% of our segmentations), it shows objects that occupy a much larger range of sizes relative to the images, and text is over five times more common in our objects (e.g., found in 22.4\% of our segmentations). Analysis of three modern few-shot localization algorithms demonstrates that they generalize poorly to our new dataset. The algorithms commonly struggle to locate objects with holes, very small and very large objects, and objects lacking text. To encourage a larger community to work on these unsolved challenges, we publicly share our annotated few-shot dataset at https://vizwiz.org .