MutualNet: Adaptive ConvNet via Mutual Learning From Different Model Configurations

MutualNet: Adaptive ConvNet via Mutual Learning From Different Model Configurations
复制标题

DOI:
10.1109/tpami.2021.3138389
复制
发表时间:
2021-05
影响因子:
23.6
通讯作者:
Taojiannan Yang;Sijie Zhu;Mat'ias Mendieta;Pu Wang;Ravikumar Balakrishnan;Minwoo Lee;T. Han;
Taojiannan Yang;Sijie Zhu;Mat'ias Mendieta;Pu Wang;Ravikumar Balakrishnan;Minwoo Lee;T. Han;
中科院分区:
计算机科学1区
文献类型:
--
作者:
Taojiannan Yang;Sijie Zhu;Mat'ias Mendieta;Pu Wang;Ravikumar Balakrishnan;Minwoo Lee;T. Han;

文献摘要

相似文献

大多数现有的深神经网络都是静态的,这意味着它们只能以固定的复杂性执行推断。但是资源预算在不同设备之间可能有很大差异。即使在单个设备上,负担得起的预算也可以随着不同的情况而变化,并且每项所需预算的反复培训网络将非常昂贵。因此,在这项工作中,我们提出了一种称为MutualNet的通用方法,以训练一个可以在各种资源约束的网络中运行的网络。我们的方法训练具有各种网络宽度和输入分辨率的模型配置队列。这种共同的学习方案不仅允许模型以不同的宽度分辨率配置运行,而且可以在这些配置之间传输独特的知识,从而帮助模型整体学习更强的表示。 MutualNet是一种通用培训方法,可以应用于各种网络结构(例如2D网络:Mobilenets,Resnet,3D网络:Slowfast,X3D)和各种任务(例如,图像分类,对象检测,分段和动作识别),,,并被证明可以在各种数据集上实现一致的改进。由于我们只训练模型一次,因此与独立培训多种模型相比,它也大大降低了培训成本。令人惊讶的是,如果动态资源限制不引起人们的关注,也可以使用互助网络来显着提高单个网络的性能。总而言之,互助网是静态和自适应,2D和3D网络的统一方法。代码和预训练模型可在https://github.com/taoyang1122/mutualnet上找到。
Most existing deep neural networks are static, which means they can only perform inference at a fixed complexity. But the resource budget can vary substantially across different devices. Even on a single device, the affordable budget can change with different scenarios, and repeatedly training networks for each required budget would be incredibly expensive. Therefore, in this work, we propose a general method called MutualNet to train a single network that can run at a diverse set of resource constraints. Our method trains a cohort of model configurations with various network widths and input resolutions. This mutual learning scheme not only allows the model to run at different width-resolution configurations but also transfers the unique knowledge among these configurations, helping the model to learn stronger representations overall. MutualNet is a general training methodology that can be applied to various network structures (e.g., 2D networks: MobileNets, ResNet, 3D networks: SlowFast, X3D) and various tasks (e.g., image classification, object detection, segmentation, and action recognition), and is demonstrated to achieve consistent improvements on a variety of datasets. Since we only train the model once, it also greatly reduces the training cost compared to independently training several models. Surprisingly, MutualNet can also be used to significantly boost the performance of a single network, if dynamic resource constraints are not a concern. In summary, MutualNet is a unified method for both static and adaptive, 2D and 3D networks. Code and pre-trained models are available at https://github.com/taoyang1122/MutualNet.