Using machine learning to improve automatic vectorization

Using machine learning to improve automatic vectorization
复制标题

使用机器学习改进自动矢量化

DOI:
10.1145/2086696.2086729
复制
发表时间:
2012
期刊:
ACM Trans. Archit. Code Optim.
影响因子:
--
通讯作者:
P. Sadayappan
P. Sadayappan
中科院分区:
--
文献类型:
--
作者:
Kevin Stock;L. Pouchet;P. Sadayappan

文献摘要

被引文献

相似文献

自动向量化对于提高计算密集型程序在现代处理器上的性能至关重要。然而,通过利用各种循环变换(例如,展开和堵塞、互换等)。 随着所考虑的转换集的增加,选择最有效的转换组合成为一个重大挑战:目前在向量化编译器中使用的成本模型通常无法识别最佳选择。在本文中,我们解决这个问题,使用机器学习模型来预测SIMD代码的性能。与使用程序高级功能的现有方法相比,我们开发了基于从生成的汇编代码中提取的功能的机器学习模型。这些模型在许多基准上进行离线训练,并在编译时用于区分从输入代码生成的许多可能的向量化变体。 我们通过使用机器学习模型来指导各种张量收缩内核的自动矢量化,展示了该模型的有效性,与英特尔ICC的自动矢量化代码相比,其改进范围从2倍到8倍不等。我们还评估了一些模板计算模型的有效性,并显示出良好的改善自动矢量化代码。
Automatic vectorization is critical to enhancing performance of compute-intensive programs on modern processors. However, there is much room for improvement over the auto-vectorization capabilities of current production compilers through careful vector-code synthesis that utilizes a variety of loop transformations (e.g., unroll-and-jam, interchange, etc.). As the set of transformations considered is increased, the selection of the most effective combination of transformations becomes a significant challenge: Currently used cost models in vectorizing compilers are often unable to identify the best choices. In this paper, we address this problem using machine learning models to predict the performance of SIMD codes. In contrast to existing approaches that have used high-level features of the program, we develop machine learning models based on features extracted from the generated assembly code. The models are trained offline on a number of benchmarks and used at compile-time to discriminate between numerous possible vectorized variants generated from the input code. We demonstrate the effectiveness of the machine learning model by using it to guide automatic vectorization on a variety of tensor contraction kernels, with improvements ranging from 2× to 8× over Intel ICC's auto-vectorized code. We also evaluate the effectiveness of the model on a number of stencil computations and show good improvement over auto-vectorized code.