Fully Convolutional Networks for Semantic Segmentation

Fully Convolutional Networks for Semantic Segmentation
复制标题

DOI:
10.1109/tpami.2016.2572683
复制
发表时间:
2017-04-01
影响因子:
23.6
通讯作者:
Darrell, Trevor
Darrell, Trevor
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shelhamer, Evan;Long, Jonathan;Darrell, Trevor

文献摘要

被引文献

相似文献

卷积网络是一种强大的视觉模型,可以产生层次结构的特征。我们证明了卷积网络本身,经过端到端,像素到像素的训练,在语义分割中改善了以前的最佳结果。我们的关键见解是构建“全卷积”网络,该网络接受任意大小的输入,并通过有效的推理和学习产生相应大小的输出。我们定义并详细描述了全卷积网络的空间,解释了它们在空间密集预测任务中的应用,并绘制了与先验模型的连接。我们将当代分类网络(AlexNet,VGG网络和GoogLeNet)适应于完全卷积网络,并通过微调将其学习的表示转移到分割任务中。然后,我们定义了一个跳过架构,结合语义信息从一个深,粗层与外观信息从一个浅,细层,以产生准确和详细的分割。我们的全卷积网络实现了PASCAL VOC(2012年平均IU为67.2%,相对提高了30%),NYUDv 2,SIFT Flow和PASCAL-Context的改进分割,而典型图像的推理时间为十分之一秒。
Convolutional networks are powerful visual models that yield hierarchies of features. We show that convolutional networks by themselves, trained end-to-end, pixels-to-pixels, improve on the previous best result in semantic segmentation. Our key insight is to build "fully convolutional" networks that take input of arbitrary size and produce correspondingly-sized output with efficient inference and learning. We define and detail the space of fully convolutional networks, explain their application to spatially dense prediction tasks, and draw connections to prior models. We adapt contemporary classification networks (AlexNet, the VGG net, and GoogLeNet) into fully convolutional networks and transfer their learned representations by fine-tuning to the segmentation task. We then define a skip architecture that combines semantic information from a deep, coarse layer with appearance information from a shallow, fine layer to produce accurate and detailed segmentations. Our fully convolutional networks achieve improved segmentation of PASCAL VOC (30% relative improvement to 67.2% mean IU on 2012), NYUDv2, SIFT Flow, and PASCAL-Context, while inference takes one tenth of a second for a typical image.