Parallel Programming Model for the Epiphany Many-Core Coprocessor Using Threaded MPI

Parallel Programming Model for the Epiphany Many-Core Coprocessor Using Threaded MPI
复制标题

使用线程 MPI 的 Epiphany 众核协处理器并行编程模型

DOI:
--
复制
发表时间:
2015
影响因子:
2.6
通讯作者:
D. Shires
D. Shires
中科院分区:
计算机科学3区
文献类型:
--
作者:
J. Ross;D. Richie;S. Park;D. Shires

文献摘要

被引文献

相似文献

Adapteva Epiphany众核架构包括具有最少非核功能的低功耗RISC内核的2D平铺网格片上网络(NoC)。它为整数和浮点计算以及并行可扩展性提供了高计算能效。然而,尽管有有趣的架构特性,一个令人信服的编程模型还没有提出日期。本文展示了一个有效的并行编程模型的Epiphany架构的消息传递接口(MPI)标准的基础上。使用MPI利用了Epiphany架构和传统的并行分布式串行核心集群之间的相似性。我们的方法使MPI代码执行的RISC阵列处理器上的修改很少,并实现高性能。我们报告了四种算法(密集矩阵-矩阵乘法,N体粒子交互,五点2D模板更新和2D FFT)的线程MPI实现的基准测试结果,并强调了快速内核间通信对架构的重要性。
The Adapteva Epiphany many-core architecture comprises a 2D tiled mesh Network-on-Chip (NoC) of low-power RISC cores with minimal uncore functionality. It offers high computational energy efficiency for both integer and floating point calculations as well as parallel scalability. Yet despite the interesting architectural features, a compelling programming model has not been presented to date. This paper demonstrates an efficient parallel programming model for the Epiphany architecture based on the Message Passing Interface (MPI) standard. Using MPI exploits the similarities between the Epiphany architecture and a conventional parallel distributed cluster of serial cores. Our approach enables MPI codes to execute on the RISC array processor with little modification and achieve high performance. We report benchmark results for the threaded MPI implementation of four algorithms (dense matrix-matrix multiplication, N-body particle interaction, a five-point 2D stencil update, and 2D FFT) and highlight the importance of fast inter-core communication for the architecture.