NeuroVectorizer: end-to-end vectorization with deep reinforcement learning

NeuroVectorizer: end-to-end vectorization with deep reinforcement learning
复制标题

DOI:
10.1145/3368826.3377928
复制
发表时间:
2019-09
期刊:
Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization
影响因子:
--
通讯作者:
Ameer Haj-Ali;Nesreen Ahmed;Theodore L. Willke;Sophia Shao;K. Asanović;I. Stoica
Ameer Haj-Ali;Nesreen Ahmed;Theodore L. Willke;Sophia Shao;K. Asanović;I. Stoica
中科院分区:
其他
文献类型:
--
作者:
Ameer Haj-Ali;Nesreen Ahmed;Theodore L. Willke;Sophia Shao;K. Asanović;I. Stoica

文献摘要

被引文献

相似文献

当编译器对当今的SIMD兼容体系结构进行矢量循环时,出现的主要挑战之一是确定矢量化或相互交织是否有益今天设计的是使用基于启发式的固定成本模型来对循环做出矢量化决策。或组织的组织。但是,软件工程师经常将每个循环的矢量化因素撰写。一种新的方法来处理循环矢量化和提案,一种使用深钢筋学习(RL)的端到端解决方案。依赖项和数据结构,可以学习一个可以更好地预测实际性能成本并确定最佳矢量化因素的复杂模型。我们提出的框架将基准代码作为输入,并提取循环代码学习的嵌入被用作深度RL代理的输入,该输入动态地确定了所有循环的矢量化因素。在当前使用的LLVM矢量器和循环多面体优化技术中,我们的实验显示1.29×-4.73×性能速度与基线相比比在广泛的基准测试中进行蛮力搜索更糟糕。
One of the key challenges arising when compilers vectorize loops for today’s SIMD-compatible architectures is to decide if vectorization or interleaving is beneficial. Then, the compiler has to determine the number of instructions to pack together and the interleaving level (stride). Compilers are designed today to use fixed-cost models that are based on heuristics to make vectorization decisions on loops. However, these models are unable to capture the data dependency, the computation graph, or the organization of instructions. Alternatively, software engineers often hand-write the vectorization factors of every loop. This, however, places a huge burden on them, since it requires prior experience and significantly increases the development time. In this work, we explore a novel approach for handling loop vectorization and propose an end-to-end solution using deep reinforcement learning (RL). We conjecture that deep RL can capture different instructions, dependencies, and data structures to enable learning a sophisticated model that can better predict the actual performance cost and determine the optimal vectorization factors. We develop an end-to-end framework, from code to vectorization, that integrates deep RL in the LLVM compiler. Our proposed framework takes benchmark codes as input and extracts the loop codes. These loop codes are then fed to a loop embedding generator that learns an embedding for these loops. Finally, the learned embeddings are used as input to a Deep RL agent, which dynamically determines the vectorization factors for all the loops. We further extend our framework to support random search, decision trees, supervised neural networks, and nearest-neighbor search. We evaluate our approaches against the currently used LLVM vectorizer and loop polyhedral optimization techniques. Our experiments show 1.29×−4.73× performance speedup compared to baseline and only 3% worse than the brute-force search on a wide range of benchmarks.