Peering Over the Memory Wall : Design Space and Performance Analysis of the Hybrid Memory Cube

Peering Over the Memory Wall : Design Space and Performance Analysis of the Hybrid Memory Cube
复制标题

越过内存墙:混合内存立方体的设计空间和性能分析

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
B. Jacob
B. Jacob
中科院分区:
--
文献类型:
--
作者:
P. Rosenfeld;E. Cooper;Todd Farrell;°. DaveResnick;B. Jacob

文献摘要

被引文献

相似文献

混合存储器立方体是一种新兴的主存储器技术,它利用3D制造技术的进步来创建CMOS逻辑层,DRAM芯片堆叠在顶部。逻辑层包含几个DRAM存储器控制器,这些控制器从主机处理器的高速串行链路接收请求。每个存储器控制器通过垂直硅通孔(TSV)连接到若干DRAM管芯中的若干存储器组。由于TSV形成具有短路径长度的密集互连,因此与传统DDRx存储器相比,控制器和存储体之间的数据总线可以以更高的吞吐量和更低的每比特能量运行。这项技术代表了主存设计的一个范式转变,可能会解决今天困扰主存系统的带宽、容量和功耗问题。虽然架构是新的,我们提出了几个设计参数扫描和性能表征突出一些有趣的问题和权衡的HMC架构。首先,我们讨论如何最好地选择链路和TSV带宽,以充分利用HMC的潜在吞吐量。接下来,我们分析了几种不同的立方体配置,资源受限,试图了解在选择系统中的内存控制器,DRAM芯片和内存库的数量的权衡。最后,提出一个一般的表征HMC的性能,我们比较执行HMC,板上缓冲,和四通道DDR3-1600主存储器系统显示,对于应用程序,可以利用HMC的并发性和带宽,一个单一的HMC可以减少整个系统的执行时间在一个非常积极的四通道DDR3-1600系统的两倍。
The Hybrid Memory Cube is an emerging main memory technology that leverages advances in 3D fabrication techniques to create a CMOS logic layer with DRAM dies stacked on top. The logic layer contains several DRAM memory controllers that receive requests from high speed serial links from the host processor. Each memory controller is connected to several memory banks in several DRAM dies with a vertical through-silicon via (TSV). Since the TSVs form a dense interconnect with short path lengths, the data bus between the controller and banks can be operated at higher throughput and lower energy per bit compared to traditional DDRx memory. This technology represents a paradigm shift in main memory design that could potentially solve the bandwidth, capacity, and power problems plaguing the main memory system today. While the architecture is new we present several design parameter sweeps and performance characterizations to highlight some interesting issues and trade-offs in the HMC architecture. First, we discuss how to best select link and TSV bandwidths to fully utilize the HMC’s potential throughput. Next, we analyze several different cube configurations that are resource constrained to try to understand the trade-offs in choosing the number of memory controllers, DRAM dies, and memory banks in the system. Finally, to present a general characterization of the HMC’s performance, we compare the execution of HMC, Buffer-on-Board, and quad channel DDR3-1600 main memory systems showing that for applications that can exploit HMC’s concurrency and bandwidth, a single HMC can reduce full-system execution time over an extremely aggressive quad channel DDR3-1600 system by a factor of two.