Lossy Compression for Embedded Computer Vision Systems

Lossy Compression for Embedded Computer Vision Systems
复制标题

DOI:
10.1109/access.2018.2852809
复制
发表时间:
2018-01-01
期刊:
影响因子:
3.9
通讯作者:
Goto, Satoshi
Goto, Satoshi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Guo, Li;Zhou, Dajiang;Goto, Satoshi

文献摘要

被引文献

相似文献

计算机视觉应用在嵌入式系统中迅速普及,这通常涉及在实时处理吞吐量约束下的视觉性能和能耗之间的困难权衡。近年来,基于FPGA和asic的硬件实现大大提高了视觉计算的能量效率。然而,这些实现通常涉及密集的内存流量,在系统级别上保留了很大一部分能源消耗。为了解决这个问题,我们是第一个提出有损压缩框架的研究人员,以利用输入图像的视觉性能和内存流量之间的权衡。为满足视觉系统对存储器访问模式的各种需求,设计了行到块格式转换框架。提出了一种基于差分脉冲码调制的梯度量化有损压缩算法。我们还介绍了它的硬件设计,支持高达12级1080p@60fps实时处理。对于基于VOC2007的定向梯度的可变形部件模型直方图,该框架在检测率下降0.05% ~ 0.34%的情况下,实现了49.6% ~ 60.5%的内存流量减少。对于ImageNet上的AlexNet,内存流量减少达到60.8%,分类率下降低于0.61%。与内存流量减少的功耗相比,所提出的输入图像压缩所涉及的开销小于5%。
Computer vision applications are rapidly gaining popularity in embedded systems, which typically involve a difficult tradeoff between vision performance and energy consumption under a constraint of real-time processing throughput. Recently, hardware (FPGA and ASIC-based) implementations have emerged, which significantly improves the energy efficiency of vision computation. These implementations, however, often involve intensive memory traffic that retains a significant portion of energy consumption at the system level. To address this issue, we are the first researchers to present a lossy compression framework to exploit the tradeoff between vision performance and memory traffic for input images. To meet various requirements for memory access patterns in the vision system, a line-to-block format conversion is designed for the framework. Differential pulse-code modulation-based gradient-oriented quantization is developed as the lossy compression algorithm. We also present its hardware design that supports up to 12-scale 1080p@60fps real-time processing. For histogram of oriented gradient-based deformable part models on VOC2007, the proposed framework achieves a 49.6%-60.5% memory traffic reduction at a detection rate degradation of 0.05%-0.34%. For AlexNet on ImageNet, memory traffic reduction achieves up to 60.8% with less than 0.61% classification rate degradation. Compared with the power consumption reduction from memory traffic, the overhead involved for the proposed input image compression is less than 5%.