BottleFit: Learning Compressed Representations in Deep Neural Networks for Effective and Efficient Split Computing

BottleFit: Learning Compressed Representations in Deep Neural Networks for Effective and Efficient Split Computing
复制标题

DOI:
10.1109/wowmom54355.2022.00032
复制
发表时间:
2022-01
期刊:
2022 IEEE 23rd International Symposium on a World of Wireless, Mobile and Multimedia Networks (WoWMoM)
影响因子:
--
通讯作者:
Yoshitomo Matsubara;Davide Callegaro;Sameer Singh;M. Levorato;Francesco Restuccia
Yoshitomo Matsubara;Davide Callegaro;Sameer Singh;M. Levorato;Francesco Restuccia
中科院分区:
其他
文献类型:
--
作者:
Yoshitomo Matsubara;Davide Callegaro;Sameer Singh;M. Levorato;Francesco Restuccia

文献摘要

被引文献

相似文献

虽然关键任务应用程序需要使用深度神经网络(DNN),但它们在移动设备上的持续执行会导致能源消耗的显著增加。虽然边缘卸载可以降低能耗,但通道质量、网络和边缘服务器负载的不稳定模式可能会导致系统的关键操作严重中断。另一种称为分割计算的方法在模型内生成压缩表示(称为“瓶颈”),以减少带宽使用和能源消耗。以前的工作已经提出了引入附加层的方法,从而损害了能量消耗和延迟。为此,我们提出了一个新的框架,称为BottleFit,它除了有针对性的DNN结构修改外,还包括一种新的训练策略,即使在强压缩率的情况下也能实现高精度。我们将BottleFit应用于图像分类中的前沿DNN模型,结果表明,BottleFit在ImageNet数据集上实现了77.1%的数据压缩,精度损失高达0.6%,而Spinn等最新技术的精度损失高达6%。我们通过实验测量了在NVIDIA Jetson Nano板(基于GPU)和Raspberry PI板(无GPU)上运行的图像分类应用程序的功耗和延迟。我们发现,与(w.r.t.)相比,BottleFit的功耗和延迟分别降低了49%和89%本地计算和37%和55%的w.r.t.边缘卸载。我们还将BottleFit与最先进的基于自动编码器的方法进行了比较,结果表明:(I)BottleFit在Jetson上的功耗和执行时间分别减少了54%和44%,在Raspberry PI上分别减少了40%和62%;(Ii)在移动设备上执行的头部模型的尺寸要小83倍。我们发布代码库,以确保研究结果的重现性。
Although mission-critical applications require the use of deep neural networks (DNNs), their continuous execution at mobile devices results in a significant increase in energy consumption. While edge offloading can decrease energy consumption, erratic patterns in channel quality, network and edge server load can lead to severe disruption of the system’s key operations. An alternative approach, called split computing, generates compressed representations within the model (called "bottlenecks"), to reduce bandwidth usage and energy consumption. Prior work has proposed approaches that introduce additional layers, to the detriment of energy consumption and latency. For this reason, we propose a new framework called BottleFit, which, in addition to targeted DNN architecture modifications, includes a novel training strategy to achieve high accuracy even with strong compression rates. We apply BottleFit on cutting-edge DNN models in image classification, and show that BottleFit achieves 77.1% data compression with up to 0.6% accuracy loss on ImageNet dataset, while state of the art such as SPINN loses up to 6% in accuracy. We experimentally measure the power consumption and latency of an image classification application running on an NVIDIA Jetson Nano board (GPU-based) and a Raspberry PI board (GPU-less). We show that BottleFit decreases power consumption and latency respectively by up to 49% and 89% with respect to (w.r.t.) local computing and by 37% and 55% w.r.t. edge offloading. We also compare BottleFit with state-of-the-art autoencoders-based approaches, and show that (i) BottleFit reduces power consumption and execution time respectively by up to 54% and 44% on the Jetson and 40% and 62% on Raspberry PI; (ii) the size of the head model executed on the mobile device is 83 times smaller. We publish the code repository for reproducibility of the results in this study.