Design of a Quantization-Based DNN Delta Compression Framework for Model Snapshots and Federated Learning

Design of a Quantization-Based DNN Delta Compression Framework for Model Snapshots and Federated Learning
复制标题

DOI:
10.1109/tpds.2022.3230840
复制
发表时间:
2023-03
影响因子:
5.3
通讯作者:
Haoyu Jin;Donglei Wu;Shuyu Zhang;Xiangyu Zou;Sian Jin;Dingwen Tao;Qing Liao;Wen Xia
Haoyu Jin;Donglei Wu;Shuyu Zhang;Xiangyu Zou;Sian Jin;Dingwen Tao;Qing Liao;Wen Xia
中科院分区:
计算机科学2区
文献类型:
--
作者:
Haoyu Jin;Donglei Wu;Shuyu Zhang;Xiangyu Zou;Sian Jin;Dingwen Tao;Qing Liao;Wen Xia

文献摘要

相似文献

深度神经网络(DNN)在许多领域都取得了巨大的成功。然而,大规模DNN在存储用于防止集群频繁故障的快照时也会带来存储成本,或者在联邦学习(FL)中传输DNN时会产生显著的通信开销。最近,诸如Delta-DNN和LC-检查点的几种方法旨在通过压缩DNN的两个相邻版本之间的差异来减小DNN的快照存储的大小(也称为,delta)。然而,我们观察到现有的方法,在DNN的增量压缩中应用传统的全局有损量化技术,不能充分利用数据的相似性,因为参数的值范围在层之间变化。为了充分挖掘delta模型的相似性,提高压缩比,本文提出了一种基于量化的局部敏感delta压缩方法QD-Compressor,该方法采用基于层的局部敏感量化方案和误差反馈机制。具体而言,量化器和量化比特数是基于增量参数的值分布和加权熵在层之间自适应的。为了避免量化误差降低恢复模型的性能,设计了一种替代的误差反馈机制,以动态地校正训练过程中的量化误差。在多个流行DNN和数据集上的实验表明,QD-Compressor在模型快照压缩场景中获得了比现有方法更高的7×-40×压缩比。此外,QD-Compressor实现了11×-15×的联合学习压缩场景的残差模型压缩比。
Deep neural networks (DNNs) have achieved remarkable success in many fields. However, large-scale DNNs also bring storage costs when storing snapshots for preventing clusters’ frequent failures or incur significant communication overheads when transmitting DNNs in the Federated Learning (FL). Recently, several approaches, such as Delta-DNN and LC-Checkpoint, aim to reduce the size of DNNs’ snapshot storage by compressing the difference between two neighboring versions of the DNNs (a.k.a., delta). However, we observe that existing approaches, applying traditional global lossy quantization techniques in DNN's delta compression, can not fully exploit the data similarity since the parameters’ value ranges vary among layers. To fully explore the similarity of the delta model and improve the compression ratio, we propose a quantization-based local-sensitive delta compression approach, named QD-Compressor, by developing a layer-based local-sensitive quantization scheme and error feedback mechanism. Specifically, the quantizers and number of quantization bits are adaptive among layers based on the value distribution and weighted entropy of the delta's parameters. To avoid quantization error degrading the performance of the restored model, an alternative error feedback mechanism is designed to dynamically correct the quantization error during the training process. Experiments on multiple popular DNNs and datasets show that QD-Compressor obtains a higher 7×-40× compression ratio in the model snapshot compression scenario than the state-of-the-art approaches. Additionally, QD-Compressor achieves an 11×-15× compression ratio to the residual model of the Federated Learning compression scenario.