PhraseCut: Language-Based Image Segmentation in the Wild

PhraseCut: Language-Based Image Segmentation in the Wild
复制标题

DOI:
10.1109/cvpr42600.2020.01023
复制
发表时间:
2020-06
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Chenyun Wu-;Zhe Lin;Scott D. Cohen;Trung Bui;Subhransu Maji
Chenyun Wu-;Zhe Lin;Scott D. Cohen;Trung Bui;Subhransu Maji
中科院分区:
其他
文献类型:
--
作者:
Chenyun Wu-;Zhe Lin;Scott D. Cohen;Trung Bui;Subhransu Maji

文献摘要

相似文献

我们考虑了给定自然语言短语的图像区域分割问题,并在包含77,262幅图像和354,486个短语-区域对的新数据集上进行了研究。我们的数据集是在视觉基因组数据集的基础上收集的,并使用现有的注释来生成一组具有挑战性的指代短语,其对应区域是手动注释的。我们的数据集中的短语对应于多个区域,描述了大量的对象和材料类别及其属性,如颜色、形状、部件以及与图像中其他实体的关系。我们的实验表明,我们数据集中概念的规模和多样性对现有的最先进技术构成了重大挑战。我们系统地处理这些概念的长尾性质,并提出了一种模块化方法来组合类别、属性和关系线索,其性能优于现有方法。
We consider the problem of segmenting image regions given a natural language phrase, and study it on a novel dataset of 77,262 images and 345,486 phrase-region pairs. Our dataset is collected on top of the Visual Genome dataset and uses the existing annotations to generate a challenging set of referring phrases for which the corresponding regions are manually annotated. Phrases in our dataset correspond to multiple regions and describe a large number of object and stuff categories as well as their attributes such as color, shape, parts, and relationships with other entities in the image. Our experiments show that the scale and diversity of concepts in our dataset poses significant challenges to the existing state-of-the-art. We systematically handle the long-tail nature of these concepts and present a modular approach to combine category, attribute, and relationship cues that outperforms existing approaches.