Zedwulf: Power-Performance Tradeoffs of a 32-Node Zynq SoC Cluster

Zedwulf: Power-Performance Tradeoffs of a 32-Node Zynq SoC Cluster
复制标题

Zedwulf:32 节点 Zynq SoC 集群的功耗与性能权衡

DOI:
--
复制
发表时间:
2015
期刊:
2015 IEEE 23rd Annual International Symposium on Field-Programmable Custom Computing Machines
影响因子:
--
通讯作者:
Nachiket Kapre
Nachiket Kapre
中科院分区:
--
文献类型:
--
作者:
P. Moorthy;Nachiket Kapre

文献摘要

被引文献

相似文献

具有混合架构的商用SoC将CPU与可编程的FPGA结构相结合,例如Xilinx Zynq SoC已经成为解决图形问题中不规则并行的具有竞争力的高能效平台。在本文中,我们用这些Zynq SoC芯片组成的32节点集群原型来加速面向通信的稀疏图应用,例如神经网络模拟。我们开发专门的MPI例程,专门为小消息流量的非常规加速器到加速器通信而开发。我们使用ARM处理器来处理MPI堆栈,同时将计算密集型计算卸载到FPGA。对于具有32M个节点和32M条边的图,在我们的研究中,Zewulf比其他x86多线程平台提供了最高的94MTEPS(每秒百万条遍历边)吞吐量1.2-1.4倍。在本实验中,当使用ARM+FPGA时,Zedwulf的运行效率为0.49 MTEPS/W,这比单独使用ARMv7 CPU的效率高1.2倍,并且在Intel Core i7-4770K平台的8%以内。
Commodity SoCs with hybrid architectures that combine CPUs with programmable FPGA fabric such as the Xilinx Zynq SoC have become a competitive energy-efficient platform for addressing irregular parallelism in graph problems. In this paper, we prototype a 32-node cluster composed from these Zynq SoC chips to accelerate communication-bound sparse graph-oriented applications such as neural network simulations. We develop specialized MPI routines specifically developed for irregular accelerator-to-accelerator communication of small message traffic. We use the ARM processor for handling the MPI stack while offloading compute-intensive calculations to the FPGA. For graphs with 32M nodes and 32M edges, Zedwulf delivers the highest 94 MTEPS (Million Traversed Edges Per Second)throughput over other x86 multi-threaded platforms in our study by 1.2 -- 1.4×. For this experiment, Zedwulf operates at an efficiency of 0.49 MTEPS/W when using ARM+FPGA which is1.2× better than using ARMv7 CPUs alone, and within 8% of the Intel Core i7-4770k platform.