Stencil codes on a vector length agnostic architecture

Stencil codes on a vector length agnostic architecture
复制标题

矢量长度不可知架构上的模板代码

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Miquel Moretó
Miquel Moretó
中科院分区:
--
文献类型:
--
作者:
Adrià Armejach;Helena Caminal;J. M. Cebrian;Rekai González;Chris Adeniyi;M. Valero;Marc Casas;Miquel Moretó

文献摘要

参考文献

被引文献

相似文献

数据级并行性经常被忽略或不足。通过向量/SIMD功能实现,它可以在广泛使用的技术(例如线程级并行性)上提供大量的性能改进。但是,手动矢量化是一个乏味且昂贵的过程,需要为每个特定的指令集或寄存器大小重复重复。此外,自动编译器矢量化易于代码复杂性,并且通常由于数据和控制依赖性而受到限制。为了解决这些问题,ARM最近发布了新的向量ISA,即可扩展向量扩展(SVE),即矢量长度不可知论(VLA)。 VLA启用不管物理矢量寄存器长度如何运行的二进制文件。在本文中,我们利用SVE的主要特征来实施和优化科学计算中无处不在的模板计算。我们表明,SVE可以轻松部署教科书优化,例如循环展开,循环融合,负载交易或数据重用。我们使用矢量长度的详细模拟范围从128到2,048位,表明这些优化可以导致比直线矢量化代码的性能改进,最高可达56.6%的2,048位矢量。此外,我们表明,由于算术强度的降低,某些优化会损害性能,并为编译器优化器提供了有用的见解。
Data-level parallelism is frequently ignored or underutilized. Achieved through vector/SIMD capabilities, it can provide substantial performance improvements on top of widely used techniques such as thread-level parallelism. However, manual vectorization is a tedious and costly process that needs to be repeated for each specific instruction set or register size. In addition, automatic compiler vectorization is susceptible to code complexity, and usually limited due to data and control dependencies. To address some these issues, Arm recently released a new vector ISA, the Scalable Vector Extension (SVE), which is Vector-Length Agnostic (VLA). VLA enables the generation of binary files that run regardless of the physical vector register length. In this paper we leverage the main characteristics of SVE to implement and optimize stencil computations, ubiquitous in scientific computing. We show that SVE enables easy deployment of textbook optimizations like loop unrolling, loop fusion, load trading or data reuse. Our detailed simulations using vector lengths ranging from 128 to 2,048 bits show that these optimizations can lead to performance improvements over straight-forward vectorized code of up to 56.6% for 2,048 bit vectors. In addition, we show that certain optimizations can hurt performance due to a reduction in arithmetic intensity, and provide insight useful for compiler optimizers.
DOI: 10.1016/j.jcp.2010.06.024
发表时间: 2010-10-01
影响因子: 4.1
作者:
Komatitsch, Dimitri;Erlebacher, Gordon;Michea, David
通讯作者: Michea, David