CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories

CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories
复制标题

DOI:
10.1145/3085572
复制
发表时间:
2017-07-01
影响因子:
1.6
通讯作者:
Srinivas, Vaishnav
Srinivas, Vaishnav
中科院分区:
计算机科学3区
文献类型:
--
作者:
Balasubramonian, Rajeev;Kahng, Andrew B.;Srinivas, Vaishnav

文献摘要

被引文献

相似文献

从历史上看,服务器设计人员通过从少数商品化的DDR内存产品中选择一种来选择简单的内存系统。我们已经目睹了片外存储器层次结构的重大变革,引入了许多新的存储器产品,如板上缓冲器、LRDIMM、HMC、HBM和nvm,仅举几例。考虑到太多的选择,不同的供应商将为他们的大容量内存系统采用不同的策略,通常会偏离DDR标准和/或在内存系统中集成新的功能。这些策略可能在互连和拓扑的选择上有所不同,在I/O和数据移动中消耗了很大一部分内存能量。为了说明内存互连专门化的情况,本文做了三个贡献。首先,我们设计了一个工具,仔细模拟内存系统中的I/O功率,探索设计空间,并使用户能够定义新的内存互连/拓扑类型。该工具针对SPICE模型进行了验证,并集成到流行的CACTI包的第7版中。我们对该工具的分析表明,几个设计参数对I/O功耗有重大影响。然后,我们使用该工具来帮助制作新颖的专用存储系统通道。我们介绍了一种新的中继板上芯片,它将DDR通道划分为多个级联通道。我们表明,对通道拓扑的这种简单更改可以将DDR DRAM的性能提高22%,并将DDR DRAM的成本降低高达65%。这种新架构不需要对内存进行任何更改,并且有效地支持混合DRAM/NVM系统。最后,作为一个更具破坏性的架构的例子,我们设计了一个定制的DIMM和并行总线,远离DDR3/DDR4标准。为了降低能耗和提高性能,基线数据通道被分成三个狭窄的并行通道,并且在较低的频率下工作。此外,这允许我们设计一个两层错误保护策略,以减少互连上的数据传输。这种架构的性能提高了18%,内存功耗降低了23%。级联通道和窄通道架构作为新工具的案例研究,并显示了重新组织基本内存互连的潜在好处。
Historically, server designers have opted for simple memory systems by picking one of a few commoditized DDR memory products. We are already witnessing a major upheaval in the off-chip memory hierarchy, with the introduction of many new memory products-buffer-on-board, LRDIMM, HMC, HBM, and NVMs, to name a few. Given the plethora of choices, it is expected that different vendors will adopt different strategies for their high-capacity memory systems, often deviating from DDR standards and/or integrating new functionality within memory systems. These strategies will likely differ in their choice of interconnect and topology, with a significant fraction of memory energy being dissipated in I/O and data movement. To make the case for memory interconnect specialization, this paper makes three contributions.First, we design a tool that carefully models I/O power in the memory system, explores the design space, and gives the user the ability to define new types of memory interconnects/topologies. The tool is validated against SPICE models, and is integrated into version 7 of the popular CACTI package. Our analysis with the tool shows that several design parameters have a significant impact on I/O power.We then use the tool to help craft novel specialized memory system channels. We introduce a new relay-on-board chip that partitions a DDR channel into multiple cascaded channels. We show that this simple change to the channel topology can improve performance by 22% for DDR DRAM and lower cost by up to 65% for DDR DRAM. This new architecture does not require any changes to DIMMs, and it efficiently supports hybrid DRAM/NVM systems.Finally, as an example of a more disruptive architecture, we design a custom DIMM and parallel bus that moves away from the DDR3/DDR4 standards. To reduce energy and improve performance, the baseline data channel is split into three narrow parallel channels and the on-DIMM interconnects are operated at a lower frequency. In addition, this allows us to design a two-tier error protection strategy that reduces data transfers on the interconnect. This architecture yields a performance improvement of 18% and a memory power reduction of 23%.The cascaded channel and narrow channel architectures serve as case studies for the new tool and show the potential for benefit from re-organizing basic memory interconnects.