DeeperLab: Single-Shot Image Parser

DeeperLab: Single-Shot Image Parser
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Tien-Ju Yang;Maxwell D. Collins;Yukun Zhu;Jyh-Jing Hwang;Ting Liu;Xiao Zhang;V. Sze;G. Papandreou-G.-Papa
Tien-Ju Yang;Maxwell D. Collins;Yukun Zhu;Jyh-Jing Hwang;Ting Liu;Xiao Zhang;V. Sze;G. Papandreou-G.-Papa
中科院分区:
其他
文献类型:
--
作者:
Tien-Ju Yang;Maxwell D. Collins;Yukun Zhu;Jyh-Jing Hwang;Ting Liu;Xiao Zhang;V. Sze;G. Papandreou-G.-Papa

文献摘要

被引文献

相似文献

我们提出了一种单镜头、自下而上的全图像解析方法。整幅图像分析,也被称为全景分割,它概括了对‘Stuff’类的语义分割和对‘Thing’类的实例分割的任务,为图像中的每个像素分配语义和实例标签。最近的整幅图像分析方法通常使用单独的独立模块来完成组成语义和实例分割任务,并且需要多次推理。相反,建议的DeeperLab图像解析器使用明显更简单的完全卷积方法执行整个图像解析,该方法以单次操作的方式联合处理语义和实例分割任务,从而产生一个更适合快速处理的流线型系统。对于定量评估,我们使用了基于实例的全景质量(PQ)度量和所提出的基于区域的解析覆盖(PC)度量,该度量更好地捕捉到了在填充类和较大对象实例上的图像解析质量。我们报告了在具有挑战性的Mapillary vistas数据集上的实验结果,其中我们的单一模型在GPU上获得了31.95%(VAL)/31.6%PQ(测试)和55.26%PC(VAL),在GPU上达到了3帧/秒(Fps)或接近实时的速度(在GPU上为22.6fps),但精度降低。
We present a single-shot, bottom-up approach for whole image parsing. Whole image parsing, also known as Panoptic Segmentation, generalizes the tasks of semantic segmentation for 'stuff' classes and instance segmentation for 'thing' classes, assigning both semantic and instance labels to every pixel in an image. Recent approaches to whole image parsing typically employ separate standalone modules for the constituent semantic and instance segmentation tasks and require multiple passes of inference. Instead, the proposed DeeperLab image parser performs whole image parsing with a significantly simpler, fully convolutional approach that jointly addresses the semantic and instance segmentation tasks in a single-shot manner, resulting in a streamlined system that better lends itself to fast processing. For quantitative evaluation, we use both the instance-based Panoptic Quality (PQ) metric and the proposed region-based Parsing Covering (PC) metric, which better captures the image parsing quality on 'stuff' classes and larger object instances. We report experimental results on the challenging Mapillary Vistas dataset, in which our single model achieves 31.95% (val) / 31.6% PQ (test) and 55.26% PC (val) with 3 frames per second (fps) on GPU or near real-time speed (22.6 fps on GPU) with reduced accuracy.