Glimpse: mathematical embedding of hardware specification for neural compilation

Glimpse: mathematical embedding of hardware specification for neural compilation
复制标题

DOI:
10.1145/3489517.3530590
复制
发表时间:
2022-07
期刊:
Proceedings of the 59th ACM/IEEE Design Automation Conference
影响因子:
--
通讯作者:
Byung Hoon Ahn;Sean Kinzer;H. Esmaeilzadeh
Byung Hoon Ahn;Sean Kinzer;H. Esmaeilzadeh
中科院分区:
其他
文献类型:
--
作者:
Byung Hoon Ahn;Sean Kinzer;H. Esmaeilzadeh

文献摘要

相似文献

深度神经网络(DNN)的成功及其计算强度预示着DNN硬件的寒武纪爆发。虽然硬件设计已经有了很大的进步,但为它们优化代码仍然是一个开放的挑战。最近的研究已经超越了传统的编译技术,采用了随机搜索算法路径,该算法盲目地生成用于真实硬件测量的相当随机的二进制样本来指导搜索。本文通过引入称为蓝图的GPU加速器硬件规范的数学嵌入,开辟了一个新的维度,以更好地指导搜索算法,并专注于具有更高潜力产生更高性能二进制文件的子空间。虽然已经提出了各种与硬件无关的样本高效技术,但没有一个最先进的编译器将硬件规范视为提高样本效率和搜索的提示。为了以数学方式将硬件规格嵌入到搜索中,我们设计了一个贝叶斯优化框架,名为Glimse,其中包含多个独一无二的组件。我们首先使用蓝图作为输入,在搜索空间中生成不同维度的先前分布。然后,我们设计了一个轻量级的神经采集函数,该函数考虑到了蓝图,以符合硬件规范,同时平衡了勘探和开发之间的权衡。最后,我们从蓝图中生成一个预测器集合,这些预测器集体投票拒绝无效的二进制样本。我们将Glimse与硬件无关的编译器进行比较。与使用多代GPU的AutoTVM[3]、Chameleon[2]和DGP[16]相比,Glimse的编译时间分别加快了6.73倍、1.51倍和1.92倍,同时也获得了最佳的推理延迟。
Success of Deep Neural Networks (DNNs) and their computational intensity has heralded Cambrian explosion of DNN hardware. While hardware design has advanced significantly, optimizing the code for them is still an open challenge. Recent research has moved past traditional compilation techniques and taken a stochastic search algorithmic path that blindly generates rather stochastic samples of the binaries for real hardware measurements to guide the search. This paper opens a new dimension by incorporating the mathematical embedding of the hardware specification of the GPU accelerators dubbed Blueprint to better guide the search algorithm and focus on sub-spaces that have higher potential for yielding higher performance binaries. While various sample efficient yet blind hardware-agnostic techniques have been proposed, none of the state-of-the-art compilers have considered hardware specification as hints to improve the sample efficiency and the search. To mathematically embed the hardware specifications into the search, we devise a Bayesian optimization framework called Glimpse with multiple exclusively unique components. We first use the Blueprint as an input to generate prior distributions of different dimensions in the search space. Then, we devise a light-weight neural acquisition function that takes into account the Blueprint to conform to the hardware specification while balancing the exploration-exploitation trade-off. Finally, we generate an ensemble of predictors from the Blueprint that collectively vote to reject invalid binary samples. We compare Glimpse with hardware-agnostic compilers. Comparison to AutoTVM [3], Chameleon [2], and DGP [16] with multiple generations of GPUs shows that Glimpse provides 6.73×, 1.51×, and 1.92× faster compilation time, respectively, while also achieving the best inference latency.