Auto-Vectorization of Loops on Intel 64 and Intel Xeon Phi: Analysis and Evaluation

Auto-Vectorization of Loops on Intel 64 and Intel Xeon Phi: Analysis and Evaluation
复制标题

Intel 64 和 Intel Xeon Phi 上循环的自动矢量化:分析和评估

DOI:
--
复制
发表时间:
2017
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
M. Kurnosov
M. Kurnosov
中科院分区:
--
文献类型:
--
作者:
O. Moldovanova;M. Kurnosov

文献摘要

被引文献

相似文献

本文评估了Intel 64和Intel Xeon Phi体系结构的现代优化编译器的自动矢量化功能,GCC C/C ++,GCC C/C ++,LLVM/Clang和PGI C/C ++。我们使用扩展的测试套件将矢量化编译器组成,该编译器由151个循环组成。在这项工作中,我们通过在标量和向量模式的不同数据类型中运行循环来估计加速,并确定编译器在研究中使用的循环类别无法矢量化。我们将双CPU系统(NUMA,2 x Intel Xeon E5-2620V4,Intel Broadwell Microarchittuction)与Intel Xeon Phi 3120a的合作处理师一起进行实验。
This paper evaluates auto-vectorizing capabilities of modern optimizing compilers Intel C/C++, GCC C/C++, LLVM/Clang and PGI C/C++ on Intel 64 and Intel Xeon Phi architectures. We use the Extended Test Suite for Vectorizing Compilers consisting of 151 loops. In this work, we estimate speedup by running the loops in scalar and vector modes for different data types and determine loop classes which the compilers used in the study fail to vectorize. We use the dual CPU system (NUMA, 2 x Intel Xeon E5-2620v4, Intel Broadwell microarchitecture) with the Intel Xeon Phi 3120A co-processor for our experiments.