Lightweight Run-Time Working Memory Compression for Deployment of Deep Neural Networks on Resource-Constrained MCUs

Lightweight Run-Time Working Memory Compression for Deployment of Deep Neural Networks on Resource-Constrained MCUs
复制标题

DOI:
10.1145/3394885.3439194
复制
发表时间:
2021-01
期刊:
2021 26th Asia and South Pacific Design Automation Conference (ASP-DAC)
影响因子:
--
通讯作者:
Zhepeng Wang;Yawen Wu;Zhenge Jia;Yiyu Shi;J. Hu
Zhepeng Wang;Yawen Wu;Zhenge Jia;Yiyu Shi;J. Hu
中科院分区:
其他
文献类型:
--
作者:
Zhepeng Wang;Yawen Wu;Zhenge Jia;Yiyu Shi;J. Hu

文献摘要

相似文献

这项工作旨在通过将深度神经网络(DNN)部署到资源约束的微控制器单元(MCUS)上来实现嵌入式设备的情报。除了低频率(例如1-16 MHz)和有限的存储空间(例如16KB至256KB ROM)外,最大的挑战之一是有限的RAM(例如2KB至64KB),这是保存中间功能所需的DNN的地图。大多数现有的神经网络压缩算法旨在减少DNN的模型大小,从而可以适应有限的存储空间。但是,它们不会大大减少中间特征图的大小,这被称为工作记忆,可能会超过RAM的能力。因此,即使在压缩后,DNN也可能无法在MCU中运行。为了解决此问题,这项工作提出了一种技术,以动态修剪运行时中间输出特征图的激活值,以确保它们可以适应有限的RAM。我们的实验结果表明,此方法可以显着减少DNN的工作记忆,以满足RAM大小的硬约束,同时保持令人满意的精度,而在内存和运行时延迟上的开销相对较低。
This work aims to achieve intelligence on embedded devices by deploying deep neural networks (DNNs) onto resource-constrained microcontroller units (MCUs). Apart from the low frequency (e.g., 1-16 MHz) and limited storage (e.g., 16KB to 256KB ROM), one of the largest challenges is the limited RAM (e.g., 2KB to 64KB), which is needed to save the intermediate feature maps of a DNN. Most existing neural network compression algorithms aim to reduce the model size of DNNs so that they can fit into limited storage. However, they do not reduce the size of intermediate feature maps significantly, which is referred to as working memory and might exceed the capacity of RAM. Therefore, it is possible that DNNs cannot run in MCUs even after compression. To address this problem, this work proposes a technique to dynamically prune the activation values of the intermediate output feature maps in the runtime to ensure that they can fit into limited RAM. The results of our experiments show that this method could significantly reduce the working memory of DNNs to satisfy the hard constraint of RAM size, while maintaining satisfactory accuracy with relatively low overhead on memory and run-time latency.