Exploring Time and Energy for Complex Accesses to a Hybrid Memory Cube

Exploring Time and Energy for Complex Accesses to a Hybrid Memory Cube
复制标题

探索对混合内存立方体进行复杂访问的时间和精力

DOI:
10.1145/2989081.2989099
复制
发表时间:
2016
期刊:
Proceedings of the Second International Symposium on Memory Systems
影响因子:
--
通讯作者:
U. Brüning
U. Brüning
中科院分区:
--
文献类型:
--
作者:
J. Schmidt;H. Fröning;U. Brüning

文献摘要

被引文献

相似文献

硅通孔(TSV)和三维芯片堆叠技术使DRAM和CMOS芯片层能够在单个堆叠内组合,从而产生堆叠存储器。先前与微处理器相关联的功能,例如存储器控制器,现在可以集成到存储器立方体中,从而允许对接口进行分组化,以提高性能并降低每比特的能耗。复杂的存储器网络变得可行,因为逻辑层可以包括路由功能。通过结合分组化接口使用TSV在不同管芯层之间的大量连接导致存储器访问带宽的实质性改进。然而,从应用程序的角度来看,利用这种巨大的带宽增加并不像看起来那么简单。在本文中,我们指出了多个陷阱时,访问堆栈存储器,即混合内存立方体(HMC)与公开可用的openHMC主机控制器相结合。HMC的内部架构仍然与传统的DRAM芯片有许多相似之处,如基于页的访问,但它在内部被划分为多个存储库。每个存储库包括存储器控制器和多个DRAM库。页面相当小,并且依赖于关闭页面策略。此外,读和写操作的比率具有应用程序应该知道的最佳值。对原子操作的内置支持听起来像是一个很好的卸载机会,但争用的影响不容忽视。除了探索这样的性能陷阱,我们进一步开始探索内存访问堆栈内存的能源效率。
Through-Silicon Vias (TSVs) and three-dimensional die stacking technologies are enabling a combination of DRAM and CMOS die layer within a single stack, leading to stacked memory. Functionality that was previously associated with the microprocessor, e.g. memory controllers, can now be integrated into the memory cube, allowing to packetize the interface for improved performance and reduced energy consumption per bit. Complex memory networks become feasible as the logic layer can include routing functionality. The massive amount of connectivity among the different die layers by the use of TSVs in combination with the packetized interface leads to a substantial improvement of memory access bandwidth. However, leveraging this vast bandwidth increase from an application point of view is not as simple as it seems. In this paper, we point out multiple pitfalls when accessing a stacked memory, namely the Hybrid Memory Cube (HMC) in combination with the publicly available openHMC host controller. HMCs internal architecture still has many similarities with traditional DRAM chips like the page-based access, but it is internally partitioned into multiple vaults. Each vault comprises a memory controller and multiple DRAM banks. Pages are rather small and rely on a closed-page policy. Also, the ratio of read and write operations has an optimum of which the application should be aware. The built-in support for atomic operations sounds like a great opportunity for off-loading, but the impact of contention cannot be neglected. Besides exploring such performance pitfalls, we further start exploring the energy efficiency of memory accesses to stacked memory.