DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs

DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
复制标题

DOI:
10.1109/tpami.2017.2699184
复制
发表时间:
2018-04-01
影响因子:
23.6
通讯作者:
Yuille, Alan L.
Yuille, Alan L.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen, Liang-Chieh;Papandreou, George;Yuille, Alan L.

文献摘要

被引文献

相似文献

在这项工作中,我们解决了使用深度学习进行语义图像分割的任务,并做出了三个主要贡献,实验表明这些贡献具有实质性的实际价值。首先,我们强调使用上采样滤波器的卷积,或“atrous卷积”,作为密集预测任务中的强大工具。Atrous卷积允许我们显式控制在深度卷积神经网络中计算特征响应的分辨率。它还允许我们有效地扩大过滤器的视野,以包含更大的上下文,而不增加参数的数量或计算量。其次,我们提出了一种新的空间金字塔池(ASPP)来鲁棒地分割多尺度的对象。ASPP以多个采样率和有效的视场使用过滤器探测传入的卷积特征层,从而以多个尺度捕获对象和图像上下文。第三,我们通过结合DCNN和概率图形模型的方法来改进对象边界的定位。DCNN中通常部署的最大池化和下采样的组合实现了不变性,但对定位精度有影响。我们通过将最终DCNN层的响应与完全连接的条件随机场(CRF)相结合来克服这一点,这在定性和定量上都表明可以提高定位性能。我们提出的“DeepLab”系统在PASCAL VOC-2012语义图像分割任务中设置了新的最先进的技术,在测试集中达到了79.7%的mIOU,并在其他三个数据集上取得了进步:PASCAL上下文,PASCAL人员部分和城市景观。我们所有的代码都在网上公开提供。
In this work we address the task of semantic image segmentation with Deep Learning and make three main contributions that are experimentally shown to have substantial practical merit. First, we highlight convolution with upsampled filters, or 'atrous convolution', as a powerful tool in dense prediction tasks. Atrous convolution allows us to explicitly control the resolution at which feature responses are computed within Deep Convolutional Neural Networks. It also allows us to effectively enlarge the field of view of filters to incorporate larger context without increasing the number of parameters or the amount of computation. Second, we propose atrous spatial pyramid pooling (ASPP) to robustly segment objects at multiple scales. ASPP probes an incoming convolutional feature layer with filters at multiple sampling rates and effective fields-of-views, thus capturing objects as well as image context at multiple scales. Third, we improve the localization of object boundaries by combining methods from DCNNs and probabilistic graphical models. The commonly deployed combination of max-pooling and downsampling in DCNNs achieves invariance but has a toll on localization accuracy. We overcome this by combining the responses at the final DCNN layer with a fully connected Conditional Random Field (CRF), which is shown both qualitatively and quantitatively to improve localization performance. Our proposed "DeepLab" system sets the new state-of-art at the PASCAL VOC-2012 semantic image segmentation task, reaching 79.7 percent mIOU in the test set, and advances the results on three other datasets: PASCAL-Context, PASCAL-Person-Part, and Cityscapes. All of our code is made publicly available online.