Automatic HBM Management: Models and Algorithms

Automatic HBM Management: Models and Algorithms
复制标题

DOI:
10.1145/3490148.3538570
复制
发表时间:
2022-07
期刊:
Proceedings of the 34th ACM Symposium on Parallelism in Algorithms and Architectures
影响因子:
--
通讯作者:
Daniel DeLayo;Kenny Zhang;Kunal Agrawal;M. A. Bender;Jonathan W. Berry;Rathish Das;Benjamin Moseley;C. Phillips
Daniel DeLayo;Kenny Zhang;Kunal Agrawal;M. A. Bender;Jonathan W. Berry;Rathish Das;Benjamin Moseley;C. Phillips
中科院分区:
其他
文献类型:
--
作者:
Daniel DeLayo;Kenny Zhang;Kunal Agrawal;M. A. Bender;Jonathan W. Berry;Rathish Das;Benjamin Moseley;C. Phillips

文献摘要

相似文献

一些过去和未来的超级计算机节点包含高带宽内存(HBM)。与标准 DRAM 相比,HBM 具有相似的延迟、更高的带宽和更低的容量。在本文中,我们评估自动管理高带宽内存的算法。之前的研究表明,在最坏的情况下,性能对于管理 DRAM 通道的策略极其敏感。先前的理论表明,基于优先级的方案(其中用于通道访问的 p 个线程之间存在静态严格的优先级顺序)是 O(1) 竞争的,但 FIFO 不是,并且在最坏的情况下是 Ω(p) 竞争的。对于目前在 DRAM 控制器硬件中使用 FIFO 变体的供应商来说,遵循这一理论指导将是一个颠覆性的变化。我们的目标是从理论上和经验上确定我们是否可以证明推荐投资基于优先级的 DRAM 控制器硬件是合理的。为了试验 DRAM 通道协议,我们选择了一个理论模型,针对真实硬件对其进行了验证,并实现了一个基本模拟器。我们证实了之前的模型理论结果,在内存带宽限制代码(GNU 排序和 TACO 稀疏矩阵向量积)的地址轨迹上运行模拟器时进行了参数扫描,并设计了更好的通道访问算法。
Some past and future supercomputer nodes incorporate High- Bandwidth Memory (HBM). Compared to standard DRAM, HBM has similar latency, higher bandwidth and lower capacity. In this paper, we evaluate algorithms for managing High- Bandwidth Memory automatically. Previous work suggests that, in the worst case, performance is extremely sensitive to the policy for managing the channel to DRAM. Prior theory shows that a priority-based scheme (where there is a static strict priority-order among p threads for channel access) is O(1)-competitive, but FIFO is not, and in the worst case is Ω(p) competitive. Following this theoretical guidance would be a disruptive change for vendors, who currently use FIFO variants in their DRAMcontroller hardware. Our goal is to determine theoretically and empirically whether we can justify recommending investment in priority-based DRAM controller hardware. In order to experiment with DRAM channel protocols, we chose a theoretical model, validated it against real hardware, and implemented a basic simulator. We corroborated the previous theoretical results for the model, conducted a parameter sweep while running our simulator on address traces from memory bandwidth-bound codes (GNU sort and TACO sparse matrix-vector product), and designed better channel-access algorithms.