All you need is superword-level parallelism: systematic control-flow vectorization with SLP

All you need is superword-level parallelism: systematic control-flow vectorization with SLP
复制标题

您所需要的只是超级字级并行性:使用 SLP 进行系统控制流矢量化

DOI:
--
复制
发表时间:
2022
期刊:
ACM-SIGPLAN Symposium on Programming Language Design and Implementation
影响因子:
--
通讯作者:
Saman P. Amarasinghe
Saman P. Amarasinghe
中科院分区:
--
文献类型:
--
作者:
Yishen Chen;Charith Mendis;Saman P. Amarasinghe

文献摘要

参考文献

被引文献

相似文献

超字级并行(SLP)向量化是一种经过验证的用于向量化直线代码的技术。它的工作原理是用等价的向量指令替换独立的同构指令。Larsen和Amarasinghe最初提出使用SLP向量化(连同循环展开)作为传统循环向量化的更简单、更灵活的替代方案。然而,这种取代传统循环向量化的愿景尚未实现,因为SLP向量化不能直接与控制流推理。在这项工作中,我们引入了SuperVectorization,一个新的矢量化框架,概括SLP矢量化,以揭示跨越不同的基本块和循环嵌套的并行性。通过系统地跨控制流区域(诸如基本块和循环)向量化指令的能力,我们的框架同时包含内循环、外循环和直线向量化器的角色,同时保留SLP向量化的灵活性(例如,部分向量化)。我们的评估表明,我们的向量化器的单个实例与LLVM的向量化管道(包括循环和SLP向量化器)相比具有竞争力,并且在许多情况下明显优于LLVM的向量化管道。例如,在法尔和马克的一个未优化的顺序体渲染器上,我们的矢量化器获得了3.28倍的加速,而我们测试的生产编译器都没有矢量化到其复杂的控制流结构。
Superword-level parallelism (SLP) vectorization is a proven technique for vectorizing straight-line code. It works by replacing independent, isomorphic instructions with equivalent vector instructions. Larsen and Amarasinghe originally proposed using SLP vectorization (together with loop unrolling) as a simpler, more flexible alternative to traditional loop vectorization. However, this vision of replacing traditional loop vectorization has not been realized because SLP vectorization cannot directly reason with control flow. In this work, we introduce SuperVectorization, a new vectorization framework that generalizes SLP vectorization to uncover parallelism that spans different basic blocks and loop nests. With the capability to systematically vectorize instructions across control-flow regions such as basic blocks and loops, our framework simultaneously subsumes the roles of inner-loop, outer-loop, and straight-line vectorizer while retaining the flexibility of SLP vectorization (e.g., partial vectorization). Our evaluation shows that a single instance of our vectorizer is competitive with and, in many cases, significantly better than LLVM’s vectorization pipeline, which includes both loop and SLP vectorizers. For example, on an unoptimized, sequential volume renderer from Pharr and Mark, our vectorizer gains a 3.28× speedup, whereas none of the production compilers that we tested vectorizes to its complex control-flow constructs.
限制自动矢量化:少即是多
DOI: 10.1109/pact.2015.32
发表时间: 2015
期刊: --
影响因子: --
作者:
Porpodas V
通讯作者: Porpodas V