Performance Evaluation on GPU-FPGA Accelerated Computing Considering Interconnections between Accelerators

Performance Evaluation on GPU-FPGA Accelerated Computing Considering Interconnections between Accelerators
复制标题

考虑加速器互连的GPU-FPGA加速计算性能评估

DOI:
10.1145/3535044.3535046
复制
发表时间:
2022
期刊:
The Proceedings of the 12th International Symposium on Highly Efficient Accelerators and Reconfigurable Technologies (HEART 2022)
影响因子:
--
通讯作者:
Boku Taisuke
Boku Taisuke
中科院分区:
--
文献类型:
--
作者:
Sano Yuka;Kobayashi Ryohei;Fujita Norihisa;Boku Taisuke

文献摘要

参考文献

被引文献

相似文献

图形处理单元(GPU)通常配备有HPC系统作为加速器,因为它们的高计算能力。GPU是强大的计算设备;然而,它们在采用部分较差并行性、非常规计算或频繁节点间通信的应用程序上效率低下。为了解决GPU的这些缺点,现场可编程门阵列(FPGA)已经出现在HPC领域,因为它们的可重新配置的能力,使特定于应用的流水线硬件和存储器系统的建设。有几项研究都集中在通过结合GPU和FPGA来提高整体应用程序的性能,实现这一目标的平台采用了将这两种设备托管在单个计算节点上的方法,但这种方法的必然性尚未讨论。在本研究中,我们使用一个天体物理应用程序进行定量评估,该应用程序执行辐射传输以模拟大爆炸后的早期宇宙。应用程序在配备有GPU和FPGA的计算节点上运行,并且GPU和FPGA计算内核从应用程序中的单个CPU(进程)启动。我们修改了代码,使GPU和FPGA计算内核能够从单独的消息传递接口(MPI)进程启动。每个MPI进程被分配到两个计算节点来运行应用程序,这两个计算节点分别仅配备GPU和FPGA,并将应用程序的执行性能与原始GPU-FPGA加速应用程序的执行性能进行比较。结果显示,与原始GPU-FPGA加速应用程序相比,性能下降约为2 - 3%,从而定量证明,即使两种设备安装在不同的计算节点上,这在实际使用中也是可以接受的,具体取决于应用程序的特性。
Graphic processing units (GPUs) are often equipped with HPC systems as accelerators because of their high computing capability. GPUs are powerful computing devices; however, they operate inefficiently on applications that employ partially poor parallelism, non-regular computation, or frequent inter-node communication. To address these shortcomings of GPUs, field-programmable gate arrays (FPGA) have been emerging in the HPC domain because their reconfigurable capabilities enable the construction of application-specific pipelined hardware and memory systems. Several studies have focused on improving overall application performance by combining GPUs and FPGAs, and the platforms for achieving this have adopted the approach of hosting these two devices on a single compute node; however, the inevitability of this approach has not been discussed.In this study, we evaluated it quantitatively using an astrophysics application that performs radiative transfer to simulate the early-stage universe after the Big Bang. The application runs on a compute node equipped with a GPU and an FPGA, and the GPU and FPGA computation kernels are launched from a single CPU (process) in the application. We modified the code to enable the launch of the GPU and FPGA computation kernels from separate message-passing interface (MPI) processes. Each MPI process was assigned to two compute nodes to run the application, which were equipped only with a GPU and FPGA, respectively, and the execution performance of the application was compared against that of the original GPU-FPGA accelerated application. The results revealed that the performance degradation compared to the original GPU-FPGA accelerated application was approximately 2 ∼ 3 %, thereby demonstrating quantitatively that even if both devices are mounted on different compute nodes, this is acceptable in practical use depending on the characteristics of the application.
一种新的光线追踪方案,用于高度并行架构上的 3D 漫射辐射传输
DOI: 10.1093/pasj/psv027
发表时间: 2014
期刊: arXiv: Instrumentation and Methods for Astrophysics
影响因子: --
作者:
Satoshi Tanaka;K. Yoshikawa;T. Okamoto;Ken Hasegawa
通讯作者: Ken Hasegawa
使用 OpenCL 加速 FPGA 上的空间辐射传输
DOI: 10.1145/3241793.3241799
发表时间: 2018
期刊: Proceedings of the 9th International Symposium on Highly-Efficient Accelerators and Reconfigurable Technologies
影响因子: --
作者:
N. Fujita;Ryohei Kobayashi;Y. Yamaguchi;Yuma Oobata;T. Boku;Makito Abe;K. Yoshikawa;M. Umemura
通讯作者: M. Umemura
oneAPI 环境下的 GPU 和 FPGA 多异构加速用于天体物理模拟
DOI: 10.1145/3492805.3492817
发表时间: 2022
期刊: HPCAsia2022: International Conference on High Performance Computing in Asia-Pacific Region
影响因子: --
作者:
Ryuta Kashino;Ryohei Kobayashi;Norihisa Fujita;Taisuke Boku
通讯作者: Taisuke Boku