Smoothed dilated convolutions for improved dense prediction

Smoothed dilated convolutions for improved dense prediction
复制标题

DOI:
10.1145/3219819.3219944
复制
发表时间:
2018-07
影响因子:
4.8
通讯作者:
Zhengyang Wang;Shuiwang Ji
Zhengyang Wang;Shuiwang Ji
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhengyang Wang;Shuiwang Ji

文献摘要

被引文献

相似文献

扩张卷积(Dilated convolutions),也被称为无环卷积(atrous convolutions),已经在深度卷积神经网络(DCNN)中被广泛研究,用于各种密集预测任务。然而,扩张卷积遭受网格伪影,这妨碍了性能。在这项工作中,我们提出了两个简单而有效的degridding方法,通过研究分解的扩张卷积。与现有的模型不同,这些模型通过关注级联的膨胀卷积层来探索解决方案,我们的方法通过平滑膨胀卷积本身来解决网格伪影。此外,我们指出,这两个degridding方法是内在相关的,并定义可分离和共享(SS)的操作,推广所提出的方法。我们进一步探索了SS操作,并提出了SS输出层,它能够通过仅替换输出层来平滑整个DCNN。我们彻底评估我们的去网格化方法和SS输出层,并通过有效的感受野分析可视化平滑效果。结果表明,我们的方法degridding在密集预测任务的性能上得到了一致的改进,同时增加了可忽略不计的额外训练参数。SS输出层的性能提高了3.3%,仅包含原输出层的9%的训练参数。
Dilated convolutions, also known as atrous convolutions, have been widely explored in deep convolutional neural networks (DCNNs) for various dense prediction tasks. However, dilated convolutions suffer from the gridding artifacts, which hampers the performance. In this work, we propose two simple yet effective degridding methods by studying a decomposition of dilated convolutions. Unlike existing models, which explore solutions by focusing on a block of cascaded dilated convolutional layers, our methods address the gridding artifacts by smoothing the dilated convolution itself. In addition, we point out that the two degridding approaches are intrinsically related and define separable and shared (SS) operations, which generalize the proposed methods. We further explore SS operations in view of operations on graphs and propose the SS output layer, which is able to smooth the entire DCNNs by only replacing the output layer. We evaluate our degridding methods and the SS output layer thoroughly, and visualize the smoothing effect through effective receptive field analysis. Results show that our methods degridding yield consistent improvements on the performance of dense prediction tasks, while adding negligible amounts of extra training parameters. And the SS output layer improves the performance by 3.3% and contains only 9% training parameters of the original output layer.