Performance Evaluation of OpenCL-Enabled Inter-FPGA Optical Link Communication Framework CIRCUS and SMI

Performance Evaluation of OpenCL-Enabled Inter-FPGA Optical Link Communication Framework CIRCUS and SMI
复制标题

支持 OpenCL 的 FPGA 间光链路通信框架 CIRCUS 和 SMI 的性能评估

DOI:
10.1145/3432261.3432266
复制
发表时间:
2021
期刊:
The International Conference on High Performance Computing in Asia-Pacific Region
影响因子:
--
通讯作者:
T. Boku
T. Boku
中科院分区:
--
文献类型:
--
作者:
Ryuta Kashino;Ryohei Kobayashi;N. Fujita;T. Boku

文献摘要

参考文献

被引文献

相似文献

近年来,现场可编程门阵列(现场可编程门阵列)作为高性能计算领域的加速器备受关注。目前的现场可编程门阵列器件的一大特点是能够通过直接光链路实现高带宽的通信性能,从而构建多个现场可编程门阵列平台,并且具有可调整性。然而,在用户应用程序上执行FPGA编程并非易事。通过更友好的编程环境,可以将现场可编程门阵列应用于多个FPGA平台上的各种高性能计算应用。在利用现场可编程门阵列通信功能实现高层次综合的几个研究中,我们重点介绍了两个系统:通信集成可识别计算系统(CIRCUS)和流消息接口(SMI),它们都可以在Intel现场可编程门阵列上实现,其直接光链路的峰值性能为40∼100Gbps。在这两个系统中,用户都可以访问OpenCL内核中的光链路,其中可以对HPC应用进行高级编程。在本文中,我们介绍了它们的实际案例,并比较了它们的实现和在实际系统中的性能。综上所述,我们使用OpenCL码对单点对点通信的Circus系统进行了评估,在100Gbps的光链路上实现了高达90Gbps的带宽。实验结果表明,在相同的平台上实现的四个FPGA之间的广播数据传输速度是SMI系统的2.7倍,达到了31Gbps的带宽,是SMI系统的5.3倍。此外,我们确定了SMI在100Gbps平台上应用时性能瓶颈的主要原因,并将其与Circus实现进行了比较。
In recent years, Field Programmable Gate Array (FPGAs) have attracted much attention as accelerators in the research area of HighPerformance Computing (HPC). One of the strong features of current FPGA devices is their ability to achieve high-bandwidth communication performance with direct optical links to construct multi-FPGA platforms as well as their adjustability. However, FPGA programming is not easily performed on user applications. By more user-friendly programming environments, FPGAs can be applied to various HPC applications on multi-FPGA platforms. Of the several studies aimed at realizing high-level synthesis to utilize the FPGA communication feature, we focus on two systems: Communication Integrated Recongurable CompUting System (CIRCUS) and Streaming Message Interface (SMI) which are available on an Intel FPGA with direct optical links with a peak performance of 40 ∼ 100 Gbps. In both systems, a user can access the optical link in OpenCL kernels where high-level programming for HPC applications is possible. In this paper, we introduce them for practical cases and compare their implementations and performance in real systems. In conclusion, we evaluated that the CIRCUS system for single point-to-point communication achieves a bandwidth of up to 90 Gbps with a 100-Gbps optical link using OpenCL code. It is 2.7 times faster than the SMI system implemented on the same platform, and we also confirmed that the broadcast data transfer among four FPGAs using CIRCUS is up to 31 Gbps of bandwidth which is 5.3 times faster compared to that achieved using SMI. In addition, we determined the main cause of the performance bottleneck on SMI when it is applied to a 100-Gbps platform and compared it with the CIRCUS implementation.
DOI: 10.1145/2678373.2665678
发表时间: 2014-10
期刊: 2014 ACM/IEEE 41st International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger
通讯作者: Andrew Putnam;Adrian M. Caulfield;Eric S. Chung;Derek Chiou;Kypros Constantinides;J. Demme;H. Esmaeilzadeh;J. Fowers;Gopi Prashanth Gopal;J. Gray;M. Haselman;S. Hauck;Stephen Heil;Amir Hormati;Joo-Young Kim;S. Lanka;J. Larus;Eric Peterson;Simon Pope;Aaron Smith;J. Thong;Phillip Yi Xiao;D. Burger