Breaking SIMD shackles with an exposed flexible microarchitecture and the access execute PDG

Breaking SIMD shackles with an exposed flexible microarchitecture and the access execute PDG
复制标题

通过暴露的灵活微架构和执行 PDG 的访问来打破 SIMD 束缚

DOI:
10.1109/pact.2013.6618830
复制
发表时间:
2013
期刊:
Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Karthikeyan Sankaralingam
Karthikeyan Sankaralingam
中科院分区:
--
文献类型:
--
作者:
Venkatraman Govindaraju;Tony Nowatzki;Karthikeyan Sankaralingam

文献摘要

被引文献

相似文献

现代的微处理器通过核心数据平行的加速器利用数据水平,以短矢量ISA的形式(例如SSE/AVX和NEON)的形式,尽管这些ISA范围已经存在数十年了,但编译器并未产生高质量的质量,没有重要的程序员干预和手动优化的代码。编译器的角色并简单地限制了编译器可以正确映射到这些数据并行加速器的代码。编译器的加速器接口实现了一个更强大,更有效的系统,支持许多类型的代码。与手动优化的实现相匹配,以解决灵活加速器的编译挑战,我们提出了一个名为“访问执行程序依赖图”的程序依赖图的变体使用此表示形式,并通过考虑为SSE开发和调整的一组内核,以及“挑战”数据并行应用程序,即Parboil基准测试。我们的编译器针对Dyser Accelerator,为内核和完整应用提供了高质量的代码,通常在手动优化的30%以内,并且表现出色的编译器生产的SSE代码。
Modern microprocessors exploit data level parallelism through in-core data-parallel accelerators in the form of short vector ISA extentions such as SSE/AVX and NEON. Although these ISA extentions have existed for decades, compilers do not generate good quality, high-performance vectorized code without significant programmer intervention and manual optimization. The fundamental problem is that the architecture is too rigid, which overly complicates the compiler's role and simultaneously restricts the types of codes that the compiler can profitably map to these data-parallel accelerators. We take a fundamentally new approach that first makes the architecture more flexible and exposes this flexibility to the compiler. Counter-intuitively, increasing the complexity of the accelerator's interface to the compiler enables a more robust and efficient system that supports many types of codes. This system also enables the performance of auto-acceleration to be comparable to that of manually-optimized implementations. To address the challenges of compiling for flexible accelerators, we propose a variant of Program Dependence Graph called the Access Execute Program Dependence Graph to capture spatio-temporal aspects of memory accesses and computations. We implement a compiler that uses this representation and evaluate it by considering both a suite of kernels developed and tuned for SSE, and “challenge” data-parallel applications, the Parboil benchmarks. We show that our compiler, which targets the DySER accelerator, provides high-quality code for the kernels and full applications, commonly reaching within 30% of manually-optimized and out-performs compiler-produced SSE code by 1.8×.