Neural Network Compression for Noisy Storage Devices

Neural Network Compression for Noisy Storage Devices
复制标题

DOI:
10.1145/3588436
复制
发表时间:
2021-02
影响因子:
2
通讯作者:
Berivan Isik;Kristy Choi;Xin Zheng;T. Weissman;Stefano Ermon;H. P. Wong;Armin Alaghi
Berivan Isik;Kristy Choi;Xin Zheng;T. Weissman;Stefano Ermon;H. P. Wong;Armin Alaghi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Berivan Isik;Kristy Choi;Xin Zheng;T. Weissman;Stefano Ermon;H. P. Wong;Armin Alaghi

文献摘要

相似文献

神经网络参数的压缩和有效存储对于运行在资源受限设备上的应用程序至关重要。尽管在神经网络模型压缩方面取得了重大进展,但在神经网络参数的实际物理存储方面的研究却相当少。通常,模型压缩和物理存储是解耦的,因为带有纠错码(ecc)的数字存储介质提供了健壮的无错误存储。然而,这种解耦方法是低效的,因为它忽略了大多数神经网络中存在的过度参数化,并强制内存设备为每个信息分配相同数量的资源,而不管其重要性。在这项工作中,我们研究了模拟存储设备作为数字媒体的替代品-它自然地提供了一种方法,为不同于其对应的重要位增加更多的保护,但如果天真地使用,它会产生噪声,并且可能会损害存储模型的性能。我们为模拟设备上的神经网络权值存储开发了多种鲁棒编码策略,并提出了一种联合优化模型压缩和内存资源分配的方法。然后,我们展示了我们的方法在现有压缩技术的MNIST、CIFAR-10和ImageNet数据集上训练的模型上的有效性。与传统的无差错数字存储相比,我们的方法将内存占用减少了一个数量级,而不会显著影响存储模型的准确性。
Compression and efficient storage of neural network (NN) parameters is critical for applications that run on resource-constrained devices. Despite the significant progress in NN model compression, there has been considerably less investigation in the actual physical storage of NN parameters. Conventionally, model compression and physical storage are decoupled, as digital storage media with error-correcting codes (ECCs) provide robust error-free storage. However, this decoupled approach is inefficient as it ignores the overparameterization present in most NNs and forces the memory device to allocate the same amount of resources to every bit of information regardless of its importance. In this work, we investigate analog memory devices as an alternative to digital media – one that naturally provides a way to add more protection for significant bits unlike its counterpart, but is noisy and may compromise the stored model’s performance if used naively. We develop a variety of robust coding strategies for NN weight storage on analog devices, and propose an approach to jointly optimize model compression and memory resource allocation. We then demonstrate the efficacy of our approach on models trained on MNIST, CIFAR-10, and ImageNet datasets for existing compression techniques. Compared to conventional error-free digital storage, our method reduces the memory footprint by up to one order of magnitude, without significantly compromising the stored model’s accuracy.