End-to-End Image Classification and Compression With Variational Autoencoders

End-to-End Image Classification and Compression With Variational Autoencoders
复制标题

DOI:
10.1109/jiot.2022.3182313
复制
发表时间:
2022-11
影响因子:
10.6
通讯作者:
Lahiru D. Chamain;Siyu Qi;Zhi Ding
Lahiru D. Chamain;Siyu Qi;Zhi Ding
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lahiru D. Chamain;Siyu Qi;Zhi Ding

文献摘要

相似文献

在过去的十年里,深度学习和人工智能在广泛的应用中占据了主导地位。特别是,无线智能手机和物联网设备的海洋继续推动基于边缘/云的机器学习(ML)系统的巨大增长,包括图像/语音识别和分类。为了克服云机器学习中有限网络带宽的基础设施障碍,现有的解决方案主要依赖于传统的压缩编解码器,如JPEG,这些编解码器历来是为人类最终用户设计的,而不是机器学习算法。在有限的带宽下,传统的编解码器不一定保留ML算法的重要功能,从而导致潜在的性能下降。这项工作研究了用于网络学习任务(如图像分类)的可编程商业编解码器设置的应用驱动优化。基于变分自编码器(VAEs)的基础,我们开发了一个端到端的网络学习框架,通过联合优化编解码器和分类器,而无需在给定的数据速率(带宽)下重建图像。与标准JPEG编解码器相比,本文提出的VAE联合压缩分类框架在数据速率为0.8 bpp的情况下,对CIFAR-10和ImageNet-1k数据集的分类准确率分别提高了10%和4%以上。我们提出的基于vae的模型显示,与基线模型相比,编码器尺寸减少了65%-99%,推理速度提高了1.5 - 13.1美元,功耗节省了25%-99%。我们进一步证明了一个简单的解码器可以在不影响分类精度的情况下以足够的质量重建图像。
The past decade has witnessed the rising dominance of deep learning and artificial intelligence in a wide range of applications. In particular, the ocean of wireless smartphones and IoT devices continue to fuel the tremendous growth of edge/cloud-based machine learning (ML) systems, including image/speech recognition and classification. To overcome the infrastructural barrier of limited network bandwidth in cloud ML, existing solutions have mainly relied on traditional compression codecs such as JPEG that were historically engineered for human-end users instead of ML algorithms. Traditional codecs do not necessarily preserve features important to ML algorithms under limited bandwidth, leading to potentially inferior performance. This work investigates application-driven optimization of programmable commercial codec settings for networked learning tasks such as image classification. Based on the foundation of variational autoencoders (VAEs), we develop an end-to-end networked learning framework by jointly optimizing the codec and classifier without reconstructing images for a given data rate (bandwidth). Compared with the standard JPEG codec, the proposed VAE joint compression and classification framework achieves classification accuracy improvement by over 10% and 4%, respectively, for CIFAR-10 and ImageNet-1k data sets at data rate of 0.8 bpp. Our proposed VAE-based models show 65%–99% reductions in encoder size, $\times 1.5$ – $\times 13.1$ improvements in inference speed, and 25%–99% savings in power compared to baseline models. We further show that a simple decoder can reconstruct images with sufficient quality without compromising classification accuracy.