Streaming-Capable High-Performance Architecture of Learned Image Compression Codecs

Streaming-Capable High-Performance Architecture of Learned Image Compression Codecs
复制标题

DOI:
10.1109/icip46576.2022.9897695
复制
发表时间:
2022-08
期刊:
2022 IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Fang-Ju Lin;Heming Sun;J. Katto
Fang-Ju Lin;Heming Sun;J. Katto
中科院分区:
其他
文献类型:
--
作者:
Fang-Ju Lin;Heming Sun;J. Katto

文献摘要

相似文献

学习图像压缩允许实现最先进的精度和压缩比,但它们相对较慢的运行时性能限制了它们的使用。虽然以前对优化学习图像编解码器的尝试更多地集中在神经模型和熵编码上,但我们提出了一种替代方法来提高各种学习图像压缩模型的运行时性能。我们引入多线程流水线和优化的内存模型,使GPU和CPU的工作负载的异步执行,充分利用计算资源。我们的架构本身已经产生了出色的性能,而无需对神经模型本身进行任何更改。我们还证明,将我们的架构与之前对神经模型的调整相结合,可以进一步提高运行时性能。我们表明,我们的实现优于吞吐量和延迟相比,基线和演示我们的实现的性能,通过创建一个实时视频流编码器-解码器的示例应用程序,与编码器运行在嵌入式设备上。
Learned image compression allows achieving state-of-the-art accuracy and compression ratios, but their relatively slow runtime performance limits their usage. While previous attempts on optimizing learned image codecs focused more on the neural model and entropy coding, we present an alternative method to improving the runtime performance of various learned image compression models. We introduce multi-threaded pipelining and an optimized memory model to enable GPU and CPU workloads’ asynchronous execution, fully taking advantage of computational resources. Our architecture alone already produces excellent performance without any change to the neural model itself. We also demonstrate that combining our architecture with previous tweaks to the neural models can further improve runtime performance. We show that our implementations excel in throughput and latency compared to the baseline and demonstrate the performance of our implementations by creating a real-time video streaming encoder-decoder sample application, with the encoder running on an embedded device.