DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network

DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network
复制标题

DOI:
10.1109/cvpr52729.2023.00597
复制
发表时间:
2023-03
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Xuan Shen;Yaohua Wang;Ming Lin;Yi-Li Huang;Hao Tang;Xiuyu Sun;Yanzhi Wang
Xuan Shen;Yaohua Wang;Ming Lin;Yi-Li Huang;Hao Tang;Xiuyu Sun;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
Xuan Shen;Yaohua Wang;Ming Lin;Yi-Li Huang;Hao Tang;Xiuyu Sun;Yanzhi Wang

文献摘要

相似文献

视觉变压器(Vision Transformer, ViT)的快速发展刷新了各种视觉任务的最新性能,使传统的基于cnn的模型黯然失色。这引发了最近CNN领域的一些反击性研究,这些研究表明,经过仔细调整,纯CNN模型可以达到与ViT模型一样好的性能。虽然令人鼓舞,但设计这种高性能的CNN模型是具有挑战性的,需要网络设计的非平凡先验知识。为此,提出了一种称为深度CNN数学架构设计的新框架(Deep- mad11源代码可在https://github.com/alibaba/lightweight-neural-architecture-search获得),以有原则的方式设计高性能CNN模型。在DeepMAD中,CNN网络被建模为一个信息处理系统,其表达能力和有效性可以通过其结构参数解析表示。然后提出了约束数学规划(MP)问题来优化这些结构参数。在内存占用很小的cpu上使用现成的MP求解器可以很容易地解决MP问题。此外,DeepMAD是一个纯数学框架:在网络设计过程中不需要GPU或训练数据。在多个大规模计算机视觉基准数据集上验证了DeepMAD的优越性。特别是在ImageNet-1k上,仅使用传统的卷积层,DeepMAD在微小水平上比ConvNeXt和Swin的top-1精度高0.7%和1.5%,在小水平上比ConvNeXt和Swin高0.8%和0.9%。
The rapid advances in Vision Transformer (ViT) refresh the state-of-the-art performances in various vision tasks, overshadowing the conventional CNN-based models. This ignites a few recent striking-back research in the CNN world showing that pure CNN models can achieve as good performance as ViT models when carefully tuned. While encouraging, designing such high-performance CNN models is challenging, requiring non-trivial prior knowledge of network design. To this end, a novel framework termed Mathematical Architecture Design for Deep CNN (Deep-MAD11Source codes are available at https://github.com/alibaba/lightweight-neural-architecture-search) is proposed to design high-performance CNN models in a principled way. In DeepMAD, a CNN network is modeled as an information processing system whose expressiveness and effectiveness can be analytically formulated by their structural parameters. Then a constrained mathematical programming (MP) problem is proposed to optimize these structural parameters. The MP problem can be easily solved by off-the-shelf MP solvers on CPUs with a small memory footprint. In addition, DeepMAD is a pure mathematical framework: no GPU or training data is required during network design. The superiority of DeepMAD is validated on multiple large-scale computer vision benchmark datasets. Notably on ImageNet-1k, only using conventional convolutional layers, DeepMAD achieves 0.7% and 1.5% higher top-1 accuracy than ConvNeXt and Swin on Tiny level, and 0.8% and 0.9% higher on Small level.