Teachers Do More Than Teach: Compressing Image-to-Image Models

Teachers Do More Than Teach: Compressing Image-to-Image Models
复制标题

DOI:
10.1109/cvpr46437.2021.01339
复制
发表时间:
2021-03
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Qing Jin;Jian Ren;Oliver J. Woodford;Jiazhuo Wang;Geng Yuan;Yanzhi Wang;S. Tulyakov
Qing Jin;Jian Ren;Oliver J. Woodford;Jiazhuo Wang;Geng Yuan;Yanzhi Wang;S. Tulyakov
中科院分区:
其他
文献类型:
--
作者:
Qing Jin;Jian Ren;Oliver J. Woodford;Jiazhuo Wang;Geng Yuan;Yanzhi Wang;S. Tulyakov

文献摘要

相似文献

生成对抗网络(GAN)在生成高保真图像方面取得了巨大的成功,但是由于巨大的计算成本和庞大的内存使用,它们的效率很低。最近在压缩GAN方面的努力表明,通过牺牲图像质量或涉及耗时的搜索过程,在获得更小的生成器方面取得了显著进展。在这项工作中,我们的目标是通过引入教师网络来解决这些问题,该网络提供了一个搜索空间,在该空间中可以找到高效的网络架构,此外还可以进行知识蒸馏。首先,我们重新审视生成模型的搜索空间,将基于接收的残差块引入生成器。其次,为了达到目标的计算成本,我们提出了一个一步修剪算法,从教师模型中搜索学生架构,大大降低了搜索成本。它不需要101稀疏正则化及其相关的超参数,简化了训练过程。最后,我们建议通过一个名为全局核对齐(GKA)的索引,通过最大限度地提高教师和学生之间的特征相似性来提取知识。我们的压缩网络实现了与原始模型相似甚至更好的图像保真度(FID,mIoU),同时大大降低了计算成本,例如,MAC代码将在https://github.com/snap-research/CAT上发布。
Generative Adversarial Networks (GANs) have achieved huge success in generating high-fidelity images, however, they suffer from low efficiency due to tremendous computational cost and bulky memory usage. Recent efforts on compression GANs show noticeable progress in obtaining smaller generators by sacrificing image quality or involving a time-consuming searching process. In this work, we aim to address these issues by introducing a teacher network that provides a search space in which efficient network architectures can be found, in addition to performing knowledge distillation. First, we revisit the search space of generative models, introducing an inception-based residual block into generators. Second, to achieve target computation cost, we propose a one-step pruning algorithm that searches a student architecture from the teacher model and substantially reduces searching cost. It requires no ℓ1 sparsity regularization and its associated hyper-parameters, simplifying the training procedure. Finally, we propose to distill knowledge through maximizing feature similarity between teacher and student via an index named Global Kernel Alignment (GKA). Our compressed networks achieve similar or even better image fidelity (FID, mIoU) than the original models with much-reduced computational cost, e.g., MACs. Code will be released at https://github.com/snap-research/CAT.