GraphIt to CUDA Compiler in 2021 LOC: A Case for High-Performance DSL Implementation via Staging with BuilDSL

GraphIt to CUDA Compiler in 2021 LOC: A Case for High-Performance DSL Implementation via Staging with BuilDSL
复制标题

DOI:
10.1109/cgo53902.2022.9741280
复制
发表时间:
2022-04
期刊:
2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Ajay Brahmakshatriya;S. Amarasinghe
Ajay Brahmakshatriya;S. Amarasinghe
中科院分区:
其他
文献类型:
--
作者:
Ajay Brahmakshatriya;S. Amarasinghe

文献摘要

被引文献

相似文献

特定于领域的语言(DSL)提供了概括和专业化之间的最佳平衡,这对于获得特定领域的最佳性能至关重要。诸如Halide和Gruphit及其丰富的调度语言之类的DSL允许用户生成最适合算法和输入的实现。 DSL还为GPU,CPU和硬件加速器等不同体系结构生成代码提供了正确的抽象。 DSL编译器庞大,通常涵盖成千上万行的代码,需要前端,一些分析和转换通行证以及特定于目标的代码生成。这些实现通常需要大量的编译器知识,而域专家则不能在没有参与编译器专家的情况下进行原型DSL。使用Scala,Ocaml或C ++等高级语言的多阶段编程是一个很好的解决方案,因为它提供了简单的解决方案使用前端和自动代码生成能力。 DSL作者通常在多阶段编程语言中将其抽象作为库实施,并使用它通过提供部分输入来生成专业代码。这仅仅是因为像Graphit这样的DSL表明需要进行多个特定于域的分析和转换以获得最佳性能,因此解决了问题。当针对诸如GPU之类的诸如GPU之类的大规模平行体系结构时,必须采取特殊护理,其中负载平衡,扭曲差异,结合记忆访问起着至关重要的作用。在本文中,我们演示了如何构建端到端DSL编译器框架和一个使用C ++中的多阶段编程的图形DSL。我们展示了如何扩展分阶段类型以执行特定领域的数据流以及控制流分析和转换。我们还展示了我们生成的CUDA代码如何匹配从最新的图形DSL Graphit生成的代码的性能。我们以实现传统DSL编译器所需的代码大小的很小一部分(8.4%)实现了所有这些。
Domain-Specific Languages (DSLs) provide the optimum balance between generalization and specialization that is crucial to getting the best performance for a particular domain. DSLs like Halide and GraphIt and their rich scheduling languages allow users to generate an implementation best suited for the algorithm and input. DSLs also provide the right abstraction for generating code for diverse architectures like GPUs, CPUs, and hardware accelerators. DSL compilers are massive, typically spanning tens of thousands of lines of code and need a frontend, some analysis and transformation passes, and target-specific code generation. These implementations usually require a great deal of compiler knowledge and domain experts cannot prototype DSLs without getting compiler experts involved.Using multi-stage programming in a high-level language like Scala, OCaml, or C++, is a great solution because it provides easy-to-use frontend and automatic code generation abilities. The DSL writers typically implement their abstraction as a library in the multi-stage programming language and use it to generate specialized code by providing partial inputs. This solves the problem only partially because DSLs like GraphIt have shown that several domain-specific analyses and transformations need to be performed to get the best performance. Special care has to be taken when targeting massively parallel architectures like GPUs where factors like load balancing, warp divergence, coalesced memory accesses play a critical role.In this paper, we demonstrate how to build an end-to-end DSL compiler framework and a graph DSL using multi-stage programming in C++. We show how the staged types can be extended to perform domain-specific data flow and control flow analyses and transformations. We also show how our generated CUDA code matches the performance of the code generated from the state-of-the-art graph DSL, GraphIt. We achieve all this in a very small fraction (8.4%) of the code size required to implement the traditional DSL compiler.