Automatic HBM Management: Models and Algorithms
Automatic HBM Management: Models and Algorithms
复制标题
DOI:
10.1145/3490148.3538570
复制
发表时间:
2022-07
期刊:
影响因子:
--
通讯作者:
Daniel DeLayo;Kenny Zhang;Kunal Agrawal;M. A. Bender;Jonathan W. Berry;Rathish Das;Benjamin Moseley;C. Phillips
中科院分区:
文献类型:
--
作者:
Daniel DeLayo;Kenny Zhang;Kunal Agrawal;M. A. Bender;Jonathan W. Berry;Rathish Das;Benjamin Moseley;C. Phillips
Some past and future supercomputer nodes incorporate High- Bandwidth Memory (HBM). Compared to standard DRAM, HBM has similar latency, higher bandwidth and lower capacity. In this paper, we evaluate algorithms for managing High- Bandwidth Memory automatically. Previous work suggests that, in the worst case, performance is extremely sensitive to the policy for managing the channel to DRAM. Prior theory shows that a priority-based scheme (where there is a static strict priority-order among p threads for channel access) is O(1)-competitive, but FIFO is not, and in the worst case is Ω(p) competitive. Following this theoretical guidance would be a disruptive change for vendors, who currently use FIFO variants in their DRAMcontroller hardware. Our goal is to determine theoretically and empirically whether we can justify recommending investment in priority-based DRAM controller hardware. In order to experiment with DRAM channel protocols, we chose a theoretical model, validated it against real hardware, and implemented a basic simulator. We corroborated the previous theoretical results for the model, conducted a parameter sweep while running our simulator on address traces from memory bandwidth-bound codes (GNU sort and TACO sparse matrix-vector product), and designed better channel-access algorithms.