Cross-Loop Optimization of Arithmetic Intensity for Finite Element Local Assembly

Cross-Loop Optimization of Arithmetic Intensity for Finite Element Local Assembly
复制标题

有限元局部装配算术强度的跨循环优化

DOI:
10.1145/2687415
复制
发表时间:
2014
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
P. Kelly
P. Kelly
中科院分区:
--
文献类型:
--
作者:
F. Luporini;A. Varbanescu;Florian Rathgeber;Gheorghe;J. Ramanujam;D. Ham;P. Kelly

文献摘要

被引文献

相似文献

我们研究和系统地评估一类可组合的代码转换,提高算术强度在本地组装操作,这代表了显着的一部分,在有限元方法的执行时间。它们的性能优化确实是一个具有挑战性的问题。尽管仿射循环嵌套通常存在,但在不同问题之间变化的短途旅行计数和数学表达式的复杂性使得难以确定成功变换的最佳序列。我们的调查导致了本地汇编内核的编译器(称为COFFEE)的实现,完全集成了开发有限元方法的框架。该编译器通过引入对编译级并行性和寄存器局部性的域感知优化来操纵从特定于域的语言生成的抽象语法树。最后,它产生C代码包括向量SIMD intrinsic。使用一系列复杂度不断增加的真实有限元问题的实验表明,实现了显着的性能改善。的方法和建议的代码转换到其他领域的适用性的一般性也进行了讨论。
We study and systematically evaluate a class of composable code transformations that improve arithmetic intensity in local assembly operations, which represent a significant fraction of the execution time in finite element methods. Their performance optimization is indeed a challenging issue. Even though affine loop nests are generally present, the short trip counts and the complexity of mathematical expressions, which vary among different problems, make it hard to determine an optimal sequence of successful transformations. Our investigation has resulted in the implementation of a compiler (called COFFEE) for local assembly kernels, fully integrated with a framework for developing finite element methods. The compiler manipulates abstract syntax trees generated from a domain-specific language by introducing domain-aware optimizations for instruction-level parallelism and register locality. Eventually, it produces C code including vector SIMD intrinsics. Experiments using a range of real-world finite element problems of increasing complexity show that significant performance improvement is achieved. The generality of the approach and the applicability of the proposed code transformations to other domains is also discussed.