GCD2: A Globally Optimizing Compiler for Mapping DNNs to Mobile DSPs

GCD2: A Globally Optimizing Compiler for Mapping DNNs to Mobile DSPs
复制标题

DOI:
10.1109/micro56248.2022.00044
复制
发表时间:
2022-10
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Wei Niu;Jiexiong Guan;Xipeng Shen;Yanzhi Wang;G. Agrawal;Bin Ren
Wei Niu;Jiexiong Guan;Xipeng Shen;Yanzhi Wang;G. Agrawal;Bin Ren
中科院分区:
其他
文献类型:
--
作者:
Wei Niu;Jiexiong Guan;Xipeng Shen;Yanzhi Wang;G. Agrawal;Bin Ren

文献摘要

相似文献

更专业的芯片正在利用可用的高晶体管密度来用更复杂的指令集大规模地揭示并行性。本文介绍了一个在移动DSP芯片上支持复杂深度神经网络(DNN)工作负载的编译系统GCD2。在充分利用这种体系结构方面,我们观察到了几个挑战,涉及SIMD宽度、更复杂的SIMD/向量指令以及具有软依赖概念的VLIW流水线。GCD2包括以下贡献:1)支持使用不同的新型SIMD指令的矩阵布局格式的开发;2)与选择用于在完整DNN中实现每个运算符的最佳指令(和相关布局)相关的全局优化问题的公式和解决方案;以及3)SDA,一种考虑软相关性的指令打包算法。这些解决方案被整合到一个完整的编译系统中,该系统使用10个大型DNN模型对其他系统进行了广泛的评估。评估结果表明,GCD2的性能比两个支持移动DSP的产品级端到端DNN执行框架(TFLite和Qualcomm SNPE)快6.0倍,比现有的三个编译器(Halide、TVM和RAKE)分别快4.5倍、3.4倍和4.0倍。GCD2在支持某些DNN的实时执行方面也是独一无二的,而它的实现使两个主要的DNN首次能够在移动DSP上执行。
More specialized chips are exploiting available high transistor density to expose parallelism at a large scale with more intricate instruction sets. This paper reports on a compilation system GCD2, developed to support complex Deep Neural Network (DNN) workloads on mobile DSP chips. We observe several challenges in fully exploiting this architecture, related to SIMD width, more complex SIMD/vector instructions, and VLIW pipeline with the notion of soft dependencies. GCD2 comprises the following contributions: 1) development of matrix layout formats that support the use of different novel SIMD instructions, 2) formulation and solution of a global optimization problem related to choosing the best instruction (and associated layout) for implementation of each operator in a complete DNN, and 3) SDA, an algorithm for packing instructions with consideration for soft dependencies. These solutions are incorporated in a complete compilation system that is extensively evaluated against other systems using 10 large DNN models. Evaluation results show that GCD2 outperforms two product-level state-of-the-art end-to-end DNN execution frameworks (TFLite and Qualcomm SNPE) that support mobile DSPs by up to $ 6.0 \times$ speedup, and outperforms three established compilers (Halide, TVM, and RAKE) by up to $4.5 \times, 3.4 \times$ and $4.0 \times$ speedup, respectively. GCD2 is also unique in supporting, real-time execution of certain DNNs, while its implementation enables two major DNNs to execute on a mobile DSP for the first time.