Wider or Deeper: Revisiting the ResNet Model for Visual Recognition

Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
复制标题

DOI:
10.1016/j.patcog.2019.01.006
复制
发表时间:
2019-06-01
影响因子:
8
通讯作者:
van den Hengel, Anton
van den Hengel, Anton
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wu, Zifeng;Shen, Chunhua;van den Hengel, Anton

文献摘要

被引文献

相似文献

社区已经在设计一个又一个前沿网络方面越来越深入,但有些作品表明我们可能在这方面走得太远了。一些研究人员将残差网络分解为指数级更宽的网络,并将残差网络的成功与大量相对较浅的模型融合在一起。由于他们早期的一些主张仍然没有得到解决,我们在本文中对这个话题进行了更多的挖掘,即,残余网络的未解视图。在此基础上,我们试图在深度和宽度之间找到一个很好的折衷。之后,我们将介绍开发基于深度学习的算法的典型流程。我们从一组相对较浅的网络开始,这些网络的性能与ImageNet分类数据集上当前(更深)的最先进模型一样好,甚至更好。然后,我们使用我们的预训练模型初始化全卷积网络(FCN),并将其调整为语义图像分割。结果表明,所提出的网络作为预训练的特征,可以大大提高现有方法的性能。即使没有用尽复杂的技术来改进经典的FCN模型,我们也可以在四个广泛使用的数据集上获得与最佳性能相当的结果,即,Cityscapes、PASCAL VOC、ADE20k和PASCAL-Context。代码和预先训练的模型被发布供公众访问。(C)2019爱思唯尔有限公司版权所有。
The community has been going deeper and deeper in designing one cutting edge network after another, yet some works are there suggesting that we may have gone too far in this dimension. Some researchers unravelled a residual network into an exponentially wider one, and assorted the success of residual networks to fusing a large amount of relatively shallow models. Since some of their early claims are still not settled, we in this paper dig more on this topic, i.e., the unravelled view of residual networks. Based on that, we try to find a good compromise between the depth and width. Afterwards, we walk through a typical pipeline of developing a deep-learning-based algorithm. We start from a group of relatively shallow networks, which perform as well or even better than the current (much deeper) state-of-the-art models on the ImageNet classification dataset. Then, we initialize fully convolutional networks (FCNs) using our pre-trained models, and tune them for semantic image segmentation. Results show that the proposed networks, as pre-trained features, can boost existing methods a lot. Even without exhausting the sophistical techniques to improve the classic FCN model, we achieve comparable results with the best performers on four widely-used datasets, i.e., Cityscapes, PASCAL VOC, ADE20k and PASCAL-Context. The code and pre-trained models are released for public access. (C) 2019 Elsevier Ltd. All rights reserved.