Hybrid memory cube performance characterization on data-centric workloads

Hybrid memory cube performance characterization on data-centric workloads
复制标题

以数据为中心的工作负载的混合内存立方体性能表征

DOI:
10.1145/2833179.2833184
复制
发表时间:
2015
期刊:
Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms
影响因子:
--
通讯作者:
C. Macaraeg
C. Macaraeg
中科院分区:
--
文献类型:
--
作者:
M. Gokhale;Scott Lloyd;C. Macaraeg

文献摘要

被引文献

相似文献

混合Memory Cube是一款早期商用产品,体现了未来堆叠式DRAM架构的属性,即大容量、高带宽、封装内存控制器和高速串行接口。我们通过在HMC FPGA板上结合仿真和执行来研究第二代HMC在以数据为中心的工作负载下的性能和能量。已使用内部的FPGA仿真器为一小部分以数据为中心的基准测试获取内存跟踪。我们的FPGA仿真器基于32位ARM处理器,非侵入性地捕获完整的内存访问轨迹,速度仅为实时的20倍。我们开发了在HMC板上运行来自多个基准的组合跟踪片段的工具,从而提供了一种独特的能力来表征数据并行工作负载下的HMC性能和功耗。我们发现,以读为主的以数据为中心的工作负载并没有很好地利用HMC的单独读写通道。我们的基准测试在HMC上实现了66%-80%的峰值带宽(对于读/写混合为50%-50%的32字节数据包,为80 GB/S),这表明在这些访问模式下,组合的读/写通道可能会显示出更高的利用率。随着对具有许多独立内存请求的高并发应用程序工作负载的需求增加,带宽线性扩展至饱和。延迟也相应增加,从极轻负载的80 ns到高带宽的130 ns不等。
The Hybrid Memory Cube is an early commercial product embodying attributes of future stacked DRAM architectures, namely large capacity, high bandwidth, on-package memory controller, and high speed serial interface. We study the performance and energy of a Gen2 HMC on data-centric workloads through a combination of emulation and execution on an HMC FPGA board. An in-house FPGA emulator has been used to obtain memory traces for a small collection of data-centric benchmarks. Our FPGA emulator is based on a 32-bit ARM processor and non-intrusively captures complete memory access traces at only 20X slowdown from real time. We have developed tools to run combined trace fragments from multiple benchmarks on the HMC board, giving a unique capability to characterize HMC performance and power usage under a data parallel workload. We find that the HMC's separate read and write channels are not well exploited by read-dominated data-centric workloads. Our benchmarks achieve between 66% -- 80% of peak bandwidth (80 GB/s for 32-byte packets with 50--50 read/write mix) on the HMC, suggesting that combined read/write channels might show higher utilization on these access patterns. Bandwidth scales linearly up to saturation with increased demand on highly concurrent application workloads with many independent memory requests. There is a corresponding increase in latency, ranging from 80 ns on an extremely light load to 130 ns at high bandwidth.