Fast HBM Access with FPGAs: Analysis, Architectures, and Applications
Fast HBM Access with FPGAs: Analysis, Architectures, and Applications
复制标题
使用 FPGA 进行快速 HBM 访问:分析、架构和应用
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
M. Reichenbach
中科院分区:
文献类型:
--
作者:
Philipp Holzinger;Daniel Reiser;Tobias Hahn;M. Reichenbach
Over the past few decades, the gap between rapidly increasing computational power and almost stagnating memory bandwidth has steadily worsened. Recently, 3D die-stacking in form of High Bandwidth Memory (HBM) enabled the first major jump in external memory throughput in years. In contrast to traditional DRAM it compensates its lower clock frequency with wide busses and a high number of separate channels. However, this also requires data to be spread out over all channels to reach the full throughput. Previous research relied on manual HBM data partitioning schemes and handled each channel as an entirely independent entity. This paper in contrast also considers scalable hardware adaptions and approaches system design holistically. In this process we first analyze the problem with real world measurements on a Xilinx HBM FPGA. Then we derive several architectural changes to improve throughput and ease accelerator design. Finally, a Roofline based model to more accurately estimate the expected performance in advance is presented. With these measures we were able to increase the throughput by up to 3.78× with random and 40.6× with certain strided access patterns compared to Xilinx’ state-of-the-art switch fabric.