UNIT: Unifying Tensorized Instruction Compilation

UNIT: Unifying Tensorized Instruction Compilation
复制标题

DOI:
10.1109/cgo51591.2021.9370330
复制
发表时间:
2021-01
期刊:
2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Jian Weng;Animesh Jain;Jie Wang;Leyuan Wang-;Yida Wang-;Tony Nowatzki
Jian Weng;Animesh Jain;Jie Wang;Leyuan Wang-;Yida Wang-;Tony Nowatzki
中科院分区:
其他
文献类型:
--
作者:
Jian Weng;Animesh Jain;Jie Wang;Leyuan Wang-;Yida Wang-;Tony Nowatzki

文献摘要

被引文献

相似文献

由于深度神经网络对密集计算的需求不断增加,研究人员开发了硬件和软件机制来减少计算和内存负担。一种广泛采用的方法是使用混合精度数据类型。然而,由于数据转换的开销,如果没有硬件专业化,就很难从混合精度中受益。最近,硬件供应商提供了专门用于混合精度张量运算的张量化指令,例如 Intel VNNI、Nvidia Tensor Core 和 ARM DOT。这些指令涉及一种新的计算习惯,它将多个低精度元素减少为一个高精度元素。由于缺乏这种新兴习惯用法的编译技术,因此很难利用这些指令。在实践中,一种方法是使用供应商提供的计算密集型内核库,但这不灵活并且会妨碍进一步的优化。另一种方法是手动编写硬件内在函数,这对于程序员来说很容易出错并且很困难。一些先前的工作试图通过为每条指令创建编译器来解决这个问题。当涉及到许多张量化指令时,这需要付出过多的努力。在这项工作中,我们开发了一个编译器框架 UNIT 来统一张量化指令的编译。这种方法的关键是统一的语义抽象,它使得新指令的集成变得容易,并且分析和转换的重用成为可能。来自不同平台的张量化指令可以通过 UNIT 进行适当的编译,以获得良好的性能。给定张量化指令和张量运算,UNIT自动检测指令的适用性,转换运算的循环组织,并重写循环体以利用张量化指令。根据我们的评估,UNIT能够针对各种主流硬件平台。生成的端到端推理模型在 x86 CPU 上比 Intel oneDNN 实现了 1.3 倍的加速,在 Nvidia GPU 上比 Nvidia cuDNN 实现了 1.75 倍的加速,在 ARM CPU 上比 ARM DOT 精心调整的 TVM 解决方案实现了 1.13 倍的加速。
Because of the increasing demand for intensive computation in deep neural networks, researchers have developed both hardware and software mechanisms to reduce the compute and memory burden. A widely adopted approach is to use mixed precision data types. However, it is hard to benefit from mixed precision without hardware specialization because of the overhead of data casting. Recently, hardware vendors offer tensorized instructions specialized for mixed-precision tensor operations, such as Intel VNNI, Nvidia Tensor Core, and ARM DOT. These instructions involve a new computing idiom, which reduces multiple low precision elements into one high precision element. The lack of compilation techniques for this emerging idiom makes it hard to utilize these instructions. In practice, one approach is to use vendor-provided libraries for computationally-intensive kernels, but this is inflexible and prevents further optimizations. Another approach is to manually write hardware intrinsics, which is error-prone and difficult for programmers. Some prior works tried to address this problem by creating compilers for each instruction. This requires excessive efforts when it comes to many tensorized instructions. In this work, we develop a compiler framework, UNIT, to unify the compilation for tensorized instructions. The key to this approach is a unified semantics abstraction which makes the integration of new instructions easy, and the reuse of the analysis and transformations possible. Tensorized instructions from different platforms can be compiled via UNIT with moderate effort for favorable performance. Given a tensorized instruction and a tensor operation, UNIT automatically detects the applicability of the instruction, transforms the loop organization of the operation, and rewrites the loop body to take advantage of the tensorized instruction. According to our evaluation, UNIT is able to target various mainstream hardware platforms. The generated end-to-end inference model achieves 1.3 x speedup over Intel oneDNN on an x86 CPU, 1.75x speedup over Nvidia cuDNN on an Nvidia GPU, and 1.13x speedup over a carefully tuned TVM solution for ARM DOT on an ARM CPU.