Supervised Compression for Resource- constrained Edge Computing Systems

Supervised Compression for Resource- constrained Edge Computing Systems
复制标题

资源受限边缘计算系统的监督压缩

DOI:
10.1109/wacv51458.2022.00100
复制
发表时间:
2022
期刊:
IEEE Winter Conference on Applications of Computer Vision (IEEE WACV
影响因子:
--
通讯作者:
Levorato, M.
Levorato, M.
中科院分区:
--
文献类型:
--
作者:
Matsubara, Y;Yang, R.;Mandt, S;Levorato, M.

文献摘要

相似文献

人们对在低功耗设备上部署深度学习算法很感兴趣,包括智能手机、无人机和医疗传感器。然而,全面的深度神经网络在能量和存储方面往往过于资源密集。因此,机器学习操作的大部分通常在边缘服务器上执行,在边缘服务器上压缩和传输数据。然而,压缩数据(如图像)会导致传输与监督任务无关的信息。另一种流行的方法是在设备和服务器之间分割深层网络,同时压缩中间功能。然而,到目前为止,这种分割计算策略由于其低效的特征压缩方法而几乎没有超过上述朴素数据压缩基线。本文采用知识提取和神经图像压缩的思想,更有效地压缩中间特征表示。我们的监督压缩方法使用教师模型和学生模型,具有随机瓶颈和可学习的熵编码先验(熵学生)。我们将我们的方法与三个视觉任务中的各种神经图像和特征压缩基线进行了比较,发现它在保持更小的端到端延迟的同时实现了更好的监督率失真性能。我们还表明,学习的特征表示可以调整为多个下游任务。
There has been much interest in deploying deep learning algorithms on low-powered devices, including smartphones, drones, and medical sensors. However, full-scale deep neural networks are often too resource-intensive in terms of energy and storage. As a result, the bulk part of the machine learning operation is therefore often carried out on an edge server, where the data is compressed and transmitted. However, compressing data (such as images) leads to transmitting information irrelevant to the supervised task. Another popular approach is to split the deep network between the device and the server while compressing intermediate features. To date, however, such split computing strategies have barely outperformed the aforementioned naive data compression baselines due to their inefficient approaches to feature compression. This paper adopts ideas from knowledge distillation and neural image compression to compress intermediate feature representations more efficiently. Our supervised compression approach uses a teacher model and a student model with a stochastic bottleneck and learnable prior for entropy coding (Entropic Student). We compare our approach to various neural image and feature compression baselines in three vision tasks and found that it achieves better supervised rate-distortion performance while maintaining smaller end-to-end latency. We furthermore show that the learned feature representations can be tuned to serve multiple downstream tasks.