Efficient Selection of Vector Instructions Using Dynamic Programming

Efficient Selection of Vector Instructions Using Dynamic Programming
复制标题

使用动态规划有效选择向量指令

DOI:
--
复制
发表时间:
2010
期刊:
Micro
影响因子:
--
通讯作者:
Vivek Sarkar
Vivek Sarkar
中科院分区:
--
文献类型:
--
作者:
R. Barik;Jisheng Zhao;Vivek Sarkar

文献摘要

被引文献

相似文献

通过SIMD向量单元提高程序性能在现代处理器中非常常见,多媒体、科学和嵌入式应用中使用SSE、MMX、VSE和VSX SIMD指令就是明证。为了充分利用向量功能,编译器需要自动生成高效的向量代码。然而,大多数商业和开源编译器没有充分利用向量单元的潜力,只为最内部的简单循环生成向量代码。本文介绍了一个动态编译器后端自动向量化框架的设计与实现,该框架不仅能生成优化的向量代码,而且能与指令调度器和寄存器分配器很好地集成在一起。该框架包括一种新颖的用于直线代码的基于{em编译时高效动态规划的}向量指令选择算法,该算法通过以下方式扩展向量化的机会:(1){em标量打包}探索将多个标量变量打包成短向量的机会;(2)如果可能的话,明智地使用{em洗牌}和{em水平}向量运算;以及(3){em代数重新关联}通过代数简化来扩展向量化的机会。我们使用Jikes RVM动态编译环境报告了自动向量化对一组标准数值基准测试的影响的性能结果。我们的结果显示,在Intel Xeon处理器上,与Tonon向量化执行相比,性能提高了57.71%,编译时间略有增加,范围从0.87%到9.992\%。对英特尔Fortran编译器(IFC)v11.1在三种基准上执行的SIMD并行化的研究表明,我们的系统在所有三种情况下都实现了向量化加速,而IFC没有。最后,将我们的方法与Cite{larsen00}中的超词级并行(SLP)算法的实现进行了比较,结果表明我们的方法相对于SLP算法有高达13.78的性能提升。
Accelerating program performance via SIMD vector units is very common in modern processors, as evidenced by the use of SSE, MMX, VSE, and VSX SIMD instructions in multimedia, scientific, and embedded applications. To take full advantage of the vector capabilities, a compiler needs to generate efficient vector code automatically. However, most commercial and open-source compilers fall short of using the full potential of vector units, and only generate vector code for simple innermost loops. In this paper, we present the design and implementation of anauto-vectorization framework in the back-end of a dynamic compiler that not only generates optimized vector code but is also well integrated with the instruction scheduler and register allocator. The framework includes a novel{em compile-time efficient dynamic programming-based} vector instruction selection algorithm for straight-line code that expands opportunities for vectorization in the following ways: (1) {em scalar packing} explores opportunities of packing multiple scalar variables into short vectors, (2)judicious use of {em shuffle} and {em horizontal} vector operations, when possible, and (3) {em algebraic reassociation} expands opportunities for vectorization by algebraic simplification. We report performance results on the impact of auto-vectorization on a set of standard numerical benchmarks using the Jikes RVM dynamic compilation environment. Our results show performance improvement of up to 57.71\% on an Intel Xeon processor, compared tonon-vectorized execution, with a modest increase in compile-time in the range from 0.87\% to 9.992\%. An investigation of the SIMD parallelization performed by v11.1 of the Intel Fortran Compiler (IFC) on three benchmarks shows that our system achieves speedup with vectorization in all three cases and IFC does not. Finally, a comparison of our approach with an implementation of the Super word Level Parallelization (SLP) algorithm from~cite{larsen00}, shows that our approach yields a performance improvement of up to 13.78\% relative to SLP.