Auto-Vectorization of Loops on Intel 64 and Intel Xeon Phi: Analysis and Evaluation
Auto-Vectorization of Loops on Intel 64 and Intel Xeon Phi: Analysis and Evaluation
复制标题
Intel 64 和 Intel Xeon Phi 上循环的自动矢量化:分析和评估
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
M. Kurnosov
中科院分区:
文献类型:
--
作者:
O. Moldovanova;M. Kurnosov
This paper evaluates auto-vectorizing capabilities of modern optimizing compilers Intel C/C++, GCC C/C++, LLVM/Clang and PGI C/C++ on Intel 64 and Intel Xeon Phi architectures. We use the Extended Test Suite for Vectorizing Compilers consisting of 151 loops. In this work, we estimate speedup by running the loops in scalar and vector modes for different data types and determine loop classes which the compilers used in the study fail to vectorize. We use the dual CPU system (NUMA, 2 x Intel Xeon E5-2620v4, Intel Broadwell microarchitecture) with the Intel Xeon Phi 3120A co-processor for our experiments.