Fast and Accurate Streaming CNN Inference via Communication Compression on the Edge

Fast and Accurate Streaming CNN Inference via Communication Compression on the Edge
复制标题

通过边缘通信压缩进行快速准确的流式 CNN 推理

DOI:
--
复制
发表时间:
2020
期刊:
International Conference on Internet-of-Things Design and Implementation
影响因子:
--
通讯作者:
B. Krishnamachari
B. Krishnamachari
中科院分区:
--
文献类型:
--
作者:
Diyi Hu;B. Krishnamachari

文献摘要

被引文献

相似文献

最近,紧凑型 CNN 模型已被开发出来,可以实现边缘计算机视觉。虽然小模型尺寸减少了存储开销,轻量级层操作减轻了边缘处理器的负担,但由于设备间带宽有限且变化多端,维持高推理性能仍然具有挑战性。我们提出了一种流式推理框架,通过通信压缩同时提高吞吐量和准确性。具体来说,我们执行以下优化: 1)分区:我们分割 CNN 层,使设备实现计算负载平衡; 2)压缩:我们识别设备间通信瓶颈,并将自动编码器插入到原始CNN中以压缩数据流量; 3)调度:当带宽变化较大时,自适应选择压缩比。由于更好的通信性能,上述优化显着提高了推理吞吐量。更重要的是,精度也提高了,因为 1) 当输入图像以高速率流入时,丢失的帧更少,2) 成功进入管道的帧得到了准确的处理,因为基于 AE 的压缩导致的信息丢失可以忽略不计。我们在 Raspberry Pi 3B+ 的管道上评估 MobileNet-v2。当平均 Wi-Fi 带宽从 3 到 9 Mbps 变化时,我们的压缩技术可将准确度提高高达 32%。
Recently, compact CNN models have been developed to enable computer vision on the edge. While the small model size reduces the storage overhead and the light-weight layer operations alleviate the burden of the edge processors, it is still challenging to sustain high inference performance due to limited and varying inter-device bandwidth. We propose a streaming inference framework to simultaneously improve throughput and accuracy by communication compression. Specifically, we perform the following optimizations: 1) Partition: we split the CNN layers such that the devices achieve computation load-balance; 2) Compression: we identify inter-device communication bottlenecks and insert Auto-Encoders into the original CNN to compress data traffic; 3) Scheduling: we adaptively select the compression ratio when the variation of bandwidth is large. The above optimizations improve inference throughput significantly due to better communication performance. More importantly, accuracy also increases since 1) fewer frames are dropped when input images are streamed in at a high rate, and 2) the frames successfully entering the pipeline are processed accurately since the AE-based compression incurs negligible information loss. We evaluate MobileNet-v2 on pipeline of Raspberry Pi 3B+. Our compression techniques lead to up to 32% accuracy improvement, when average Wi-Fi bandwidth varies from 3 to 9Mbps.