Exact Memory- and Communication-aware Scheduling of DNNs on Pipelined Edge TPUs

Exact Memory- and Communication-aware Scheduling of DNNs on Pipelined Edge TPUs
复制标题

DOI:
10.1109/sec54971.2022.00023
复制
发表时间:
2022-12
期刊:
2022 IEEE/ACM 7th Symposium on Edge Computing (SEC)
影响因子:
--
通讯作者:
Jiaqi Yin;Zhiru Zhang;Cunxi Yu
Jiaqi Yin;Zhiru Zhang;Cunxi Yu
中科院分区:
其他
文献类型:
--
作者:
Jiaqi Yin;Zhiru Zhang;Cunxi Yu

文献摘要

相似文献

深神经网络(DNN)代表许多应用程序中的最新技术,但具有大量的计算和内存要求,这极大地限制了他们在现实世界中的培训和部署。特别是,部署挑战进一步增加了边缘系统,其资源受限得多(例如,计算和内存有限),最近在许多应用程序方案中引起了重大兴趣。 Edge TPU之类的设备通常提供有限的片上存储和内存带宽,由于缺乏性能保证,因此基于启发式的提前汇编技术在优化推理性能方面受到了极大的限制。这项工作提出了一个新颖的精确管道调度框架,该框架可以实现模型参数缓存,数据依赖关系以及设备与设备通信感知的多目标优化。该框架由新型多功能SDC+ILP配方提供动力,支持命题逻辑和非平等约束。实验结果表明,所提出的调度框架始终优于商业边缘TPU编译器,在物理管道的边缘TPU设置中,在11个Imagenet模型上具有多达4个X的速度。此外,我们已经证明了使用高精度功率计测量的一致的现实世界能源效率提高。最后,所提出的框架还证明了管道边缘TPU系统上多模型共同部署的能力,该系统不受边缘TPU编译器的支持。
Deep neural networks (DNNs) represent the state-of-the-art in many applications but have substantial computational and memory requirements, which greatly limit their training and deployment in real-world systems. In particular, the deployment challenges further increase on edge systems with much more restricted resource-constrained (e.g., computation and memory bounded), which recently attracted significant interest in many application scenarios. Such devices like Edge TPUs usually provide limited on-chip storage and memory bandwidth, where the heuristic-based ahead-of-time compilation techniques are highly limited in optimizing the inference performance due to the lacks of performance guarantees. This work proposes a novel exact pipeline scheduling framework that enables model parameter caching, data dependency, and device-to-device communication-aware multi-objective optimizations. The framework is powered by novel versatile SDC+ILP formulations supporting both propositional logic and non-equality constraints. The experimental results demonstrate that the proposed scheduling frameworks consistently outperform commercial Edge TPU Compiler with up to more than 4 x speedups on eleven ImageNet models in physical pipelined Edge TPU setups. In addition, we have demonstrated consistent real-world energy efficiency improvements measured with high precision power meter. Finally, the proposed framework has also demonstrated the capability in multi-model co-deployment on pipeline Edge TPU system, which is not supported by Edge TPU Compiler.