Hardware Accelerator Integration Tradeoffs for High-Performance Computing: A Case Study of GEMM Acceleration in N-Body Methods

Hardware Accelerator Integration Tradeoffs for High-Performance Computing: A Case Study of GEMM Acceleration in N-Body Methods
复制标题

高性能计算的硬件加速器集成权衡:N 体方法中 GEMM 加速的案例研究

DOI:
10.1109/tpds.2021.3056045
复制
发表时间:
2021
影响因子:
5.3
通讯作者:
Gerstlauer, Andreas
Gerstlauer, Andreas
中科院分区:
计算机科学2区
文献类型:
--
作者:
Asri, Mochamad;Malhotra, Dhairya;Wang, Jiajun;Biros, George;John, Lizy K;Gerstlauer, Andreas

文献摘要

参考文献

相似文献

在本文中,我们研究了在不同硬件配置和使用场景下,最先进的快速多极子方法(FMM)(一种流行的N体方法)的硬件加速的性能和节能效益。我们使用专用的专用集成电路(ASIC)来加速通用矩阵-矩阵乘法(GEMM)运算。FMM广泛应用于各种应用中,并且是许多HPC应用的工作负载的代表性示例。我们比较架构,集成的GEMM ASIC旁边,在或附近的主存储器与片上耦合,旨在最大限度地减少或避免重复往返传输通过DRAM之间的通信加速器和CPU。我们研究权衡使用详细和准确校准的x86 CPU,加速器和DRAM模拟。我们的研究结果表明,简单地将加速器移近芯片并不一定会带来性能/能量的提高。我们证明,虽然仔细的软件阻塞和片上布局优化可以减少DRAM访问的2倍以上的天真的片上集成,这些戏剧性的节省DRAM流量不会自动转化为显着的总能源或运行时间节省。这主要是由于现代系统的应用特点、高空闲功率和有效隐藏存储器延迟。只有当应用软件流水线和重叠等更积极的协同优化时,才能在基线加速基础上分别实现37%和35%的额外性能和节能。当类似的优化(流水线和重叠)应用于片外集成时,片内集成的性能比片外集成高出20%,总能耗比片外集成低17%。
In this article, we study performance and energy saving benefits of hardware acceleration under different hardware configurations and usage scenarios for a state-of-the-art Fast Multipole Method (FMM), which is a popular N-body method. We use a dedicated Application Specific Integrated Circuit (ASIC) to accelerate General Matrix-Matrix Multiply (GEMM) operations. FMM is widely used in applications and is representative example of the workload for many HPC applications. We compare architectures that integrate the GEMM ASIC next to, in or near main memory with an on-chip coupling aimed at minimizing or avoiding repeated round-trip transfers through DRAM for communication between accelerator and CPU. We study tradeoffs using detailed and accurately calibrated x86 CPU, accelerator and DRAM simulations. Our results show that simply moving accelerators closer to the chip does not necessarily lead to performance/energy gains. We demonstrate that, while careful software blocking and on-chip placement optimizations can reduce DRAM accesses by 2X over a naive on-chip integration, these dramatic savings in DRAM traffic do not automatically translate into significant total energy or runtime savings. This is chiefly due to the application characteristics, the high idle power and effective hiding of memory latencies in modern systems. Only when more aggressive co-optimizations such as software pipelining and overlapping are applied, additional performance and energy savings can be unlocked by 37 and 35 percent respectively over baseline acceleration. When similar optimizations (pipelining and overlapping) are applied with an off-chip integration, on-chip integration delivers up to 20 percent better performance and 17 percent less total energy consumption than off-chip integration.
多核性能的诊断、调整和重新设计:快速多极方法的案例研究
DOI: --
发表时间: 2010
期刊: 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
Aparna Chandramowlishwaran;Kamesh Madduri;R. Vuduc
通讯作者: R. Vuduc
DOI: --
发表时间: 2010
影响因子: 1.3
作者:
Calin Cascaval;S. Chatterjee;H. Franke;K. Gildea;P. Pattnaik
通讯作者: P. Pattnaik
揭示 ASIC 上 FMM 的可行性:在 FPGA 上高效实现 N 体问题
DOI: --
发表时间: 2010
期刊: IEEE International Conference on Computational Science and Engineering
影响因子: --
作者:
Zhe Zheng;Yongxin Zhu;Xu Wang;Zhiqiang Que;Tian Huang;X. Yin;Hui Wang;G. Rong;Meikang Qiu
通讯作者: Meikang Qiu
DOI: --
发表时间: 2011
期刊:
影响因子: --
作者:
Y. Chai;W. Shen;W. Xu;Yanheng Zheng
通讯作者: Yanheng Zheng
DOI: 10.1109/hpca.2017.21
发表时间: 2017-02
期刊: 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子: --
作者:
Shaizeen Aga;Supreet Jeloka;Arun K. Subramaniyan;S. Narayanasamy;D. Blaauw;R. Das
通讯作者: Shaizeen Aga;Supreet Jeloka;Arun K. Subramaniyan;S. Narayanasamy;D. Blaauw;R. Das