A programming system for xeon phis with runtime SIMD parallelization

A programming system for xeon phis with runtime SIMD parallelization
复制标题

具有运行时 SIMD 并行化的 xeon phi 编程系统

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Supercomputing
影响因子:
--
通讯作者:
G. Agrawal
G. Agrawal
中科院分区:
--
文献类型:
--
作者:
Xin Huo;Bin Ren;G. Agrawal

文献摘要

被引文献

相似文献

英特尔至强融核提供了一个很有前途的协处理解决方案,因为它是基于流行的x86指令集。然而,为了充分利用其潜力,除了有效的大规模共享内存并行性之外,还必须对应用程序进行向量化以利用宽SIMD通道。与具有CUDA或OpenCL的GPGPU上的SIMT执行模型相比,具有类似SSE指令集的SIMD并行性施加了许多限制,并且在过去通常不会使涉及分支、不规则访问甚至减少的应用受益。在本文中,我们考虑的问题,加速应用程序涉及不同的通信模式的至强Phis,重点是有效地利用现有的SIMD并行。我们提供了一个API共享内存和SIMD并行化,并演示其实现。我们使用重载函数的实现作为提供SIMD代码的机制,这是由运行时数据重新排序和我们的方法来有效地管理控制流。我们的广泛的评估与6个流行的应用程序显示了生产(ICC)编译器实现的SIMD并行化的巨大收益,我们甚至优于OpenMP的MIMD并行。
The Intel Xeon Phi offers a promising solution to coprocessing, since it is based on the popular x86 instruction set. However, to fully utilize its potential, applications must be vectorized to leverage the wide SIMD lanes, in addition to effective large-scale shared memory parallelism. Compared to the SIMT execution model on GPGPUs with CUDA or OpenCL, SIMD parallelism with a SSE-like instruction set imposes many restrictions, and has generally not benefitted applications involving branches, irregular accesses, or even reductions in the past. In this paper, we consider the problem of accelerating applications involving different communication patterns on Xeon Phis, with an emphasis on effectively using available SIMD parallelism. We offer an API for both shared memory and SIMD parallelization, and demonstrate its implementation. We use implementations of overloaded functions as a mechanism for providing SIMD code, which is assisted by runtime data reordering and our methods to effectively manage control flow. Our extensive evaluation with 6 popular applications shows large gains over the SIMD parallelization achieved by the production (ICC) compiler, and we even outperform OpenMP for MIMD parallelism.