Look-ahead SLP: auto-vectorization in the presence of commutative operations

Look-ahead SLP: auto-vectorization in the presence of commutative operations
复制标题

前瞻 SLP:存在交换运算时的自动矢量化

DOI:
--
复制
发表时间:
2018
期刊:
IEEE/ACM International Symposium on Code Generation and Optimization
影响因子:
--
通讯作者:
L. F. Góes
L. F. Góes
中科院分区:
--
文献类型:
--
作者:
Vasileios Porpodas;Rodrigo C. O. Rocha;L. F. Góes

文献摘要

参考文献

被引文献

相似文献

自动向量化编译器自动从标量代码生成向量(SIMD)指令。最先进的直线代码矢量化算法是超字级矢量化(SLP)。在这项工作中,我们确定了SLP算法的核心的主要限制,在收集矢量化候选指令,形成SLP图数据结构的性能关键步骤。SLP在构建其向量化图时缺乏全局知识,这在遇到交换指令时对其局部决策产生负面影响。我们提出了LSLP,一个改进的算法,可以插入到现有的SLP实现,并可以有效地向量化代码与任意长链的交换操作。LSLP依赖于短期深度前瞻,以便更好地了解本地决策。我们在真实的机器上的评估表明,LSLP可以显着提高真实世界的代码的性能,很少的编译时间开销。
Auto-vectorizing compilers automatically generate vector (SIMD) instructions out of scalar code. The state-of-the-art algorithm for straight-line code vectorization is Superword-Level Parallelism (SLP). In this work we identify a major limitation at the core of the SLP algorithm, in the performance-critical step of collecting the vectorization candidate instructions that form the SLP-graph data structure. SLP lacks global knowledge when building its vectorization graph, which negatively affects its local decisions when it encounters commutative instructions. We propose LSLP, an improved algorithm that can plug-in to existing SLP implementations, and can effectively vectorize code with arbitrarily long chains of commutative operations. LSLP relies on short-depth look-ahead for better-informed local decisions. Our evaluation on a real machine shows that LSLP can significantly improve the performance of real-world code with little compilation-time overhead.
限制自动矢量化:少即是多
DOI: 10.1109/pact.2015.32
发表时间: 2015
期刊: --
影响因子: --
作者:
Porpodas V
通讯作者: Porpodas V