Automated Quantization and Retraining for Neural Network Models Without Labeled Data

Automated Quantization and Retraining for Neural Network Models Without Labeled Data
复制标题

无标记数据的神经网络模型的自动量化和再训练

DOI:
10.1109/access.2022.3190627
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Hajimu Iida
Hajimu Iida
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kundjanasith Thonglek;Keichi Takahashi;Kohei Ichikawa;Chawanat Nakasan;Hidemoto Nakada;Ryousei Takano;Pattara Leelaprute;Hajimu Iida

文献摘要

相似文献

将神经网络模型部署到边缘设备越来越受欢迎,因为这样的部署减少了响应时间,并确保了更好的服务数据隐私。然而,由于计算资源和存储空间有限,在边缘设备上运行大型模型会带来挑战。因此,研究人员提出了各种模型压缩方法,以减少模型的大小。为了平衡模型大小和准确性之间的权衡,传统的模型压缩方法需要人工努力来找到在不显著降低准确性的情况下减小模型大小的最佳配置。在这篇文章中,我们提出了一种方法来自动找到量化的最佳配置。所提出的方法建议多个压缩配置,产生具有不同大小和精度的模型,用户可以从中选择适合其用例的配置。此外,我们提出了一种不需要任何标记数据集进行再训练的再训练方法。我们使用各种神经网络模型进行分类,回归和语义相似性任务评估了所提出的方法,并证明所提出的方法将模型的大小减少了至少30%,同时保持不到1%的准确性损失。我们将所提出的方法与最先进的自动压缩方法进行了比较,并表明它可以提供比现有方法更好的压缩配置。
Deploying neural network models to edge devices is becoming increasingly popular because such deployment decreases the response time and ensures better data privacy of services. However, running large models on edge devices poses challenges because of limited computing resources and storage space. Researchers have therefore proposed various model compression methods to reduce the model size. To balance the trade-off between model size and accuracy, conventional model compression methods require manual effort to find the optimal configuration that reduces the model size without significant degradation of accuracy. In this article, we propose a method to automatically find the optimal configurations for quantization. The proposed method suggests multiple compression configurations that produce models with different size and accuracy, from which users can select the configurations that suit their use cases. Additionally, we propose a retraining method that does not require any labeled datasets for retraining. We evaluated the proposed method using various neural network models for classification, regression and semantic similarity tasks, and demonstrated that the proposed method reduced the size of models by at least 30% while maintaining less than 1% loss of accuracy. We compared the proposed method with state-of-the-art automated compression methods, and showed that it can provide better compression configurations than existing methods.