OTOV2: Automatic, Generic, User-Friendly

OTOV2: Automatic, Generic, User-Friendly
复制标题

DOI:
10.48550/arxiv.2303.06862
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Tianyi Chen;Luming Liang;Tian Ding;Zhihui Zhu;Ilya Zharkov
Tianyi Chen;Luming Liang;Tian Ding;Zhihui Zhu;Ilya Zharkov
中科院分区:
其他
文献类型:
--
作者:
Tianyi Chen;Luming Liang;Tian Ding;Zhihui Zhu;Ilya Zharkov

文献摘要

相似文献

现有的基于结构剪枝的模型压缩方法通常需要复杂的多阶段过程。每个单独的阶段都需要来自最终用户的大量工程工作和领域知识,这阻碍了它们在更广泛的场景中的更广泛应用。我们提出了第二代只训练一次(OTOv2),它首先从零开始自动训练和压缩一般的DNN一次,以产生更紧凑的、具有竞争力的模型,而不需要微调。OTOv2是自动的,可插入到各种深度学习应用程序中,用户只需几乎最少的工程工作。在方法上,OTOv2提出了两个主要的改进:(I)自治性:自动利用一般DNN的依赖关系,将可训练变量划分为零不变组(ZIG),并构建压缩模型;(Ii)Dual Half-Space Projected Gate(DHSPG):一种新的优化器,可以更可靠地解决结构化稀疏问题。我们在VGG、ResNet、CARN、ConvNeXt、DenseNet和StackedUnet等多种模型体系结构上数值演示了OTOv2的通用性和自主性,其中大部分不需要大量的手工工作就不能用其他方法来处理。结合CIFAR10/100、DIV2K、Fashion-MNIST、SVNH和ImageNet等基准数据集,通过具有竞争力的表现甚至更好的表现来验证其有效性。源代码可在https://github.com/tianyic/only_train_once.上找到
The existing model compression methods via structured pruning typically require complicated multi-stage procedures. Each individual stage necessitates numerous engineering efforts and domain-knowledge from the end-users which prevent their wider applications onto broader scenarios. We propose the second generation of Only-Train-Once (OTOv2), which first automatically trains and compresses a general DNN only once from scratch to produce a more compact model with competitive performance without fine-tuning. OTOv2 is automatic and pluggable into various deep learning applications, and requires almost minimal engineering efforts from the users. Methodologically, OTOv2 proposes two major improvements: (i) Autonomy: automatically exploits the dependency of general DNNs, partitions the trainable variables into Zero-Invariant Groups (ZIGs), and constructs the compressed model; and (ii) Dual Half-Space Projected Gradient (DHSPG): a novel optimizer to more reliably solve structured-sparsity problems. Numerically, we demonstrate the generality and autonomy of OTOv2 on a variety of model architectures such as VGG, ResNet, CARN, ConvNeXt, DenseNet and StackedUnets, the majority of which cannot be handled by other methods without extensive handcrafting efforts. Together with benchmark datasets including CIFAR10/100, DIV2K, Fashion-MNIST, SVNH and ImageNet, its effectiveness is validated by performing competitively or even better than the state-of-the-arts. The source code is available at https://github.com/tianyic/only_train_once.