There and back again: Optimizing the interconnect in networks of memory cubes

There and back again: Optimizing the interconnect in networks of memory cubes
复制标题

来来回回:优化内存立方体网络中的互连

DOI:
--
复制
发表时间:
2017
期刊:
International Symposium on Computer Architecture
影响因子:
--
通讯作者:
G. Loh
G. Loh
中科院分区:
--
文献类型:
--
作者:
Matthew Poremba;Itir Akgun;Jieming Yin;Onur Kayiran;Yuan Xie;G. Loh

文献摘要

被引文献

相似文献

高性能计算、企业和数据中心服务器正在推动对更高的总内存容量和内存性能的需求。具有高每个封装容量(来自3D集成)的内存“立方体”以及高速点对点互连提供了可扩展的内存系统架构,具有提供容量和性能的潜力。多个这样的多维数据集连接在一起可以形成一个“内存网络”(Memory Network, MN),但是这种MN的设计空间非常大,每个内存多维数据集包括多种拓扑类型和多种内存技术。在这项工作中,我们首先分析了几种具有不同内存封装技术混合的MN拓扑,以了解此类系统的关键权衡和瓶颈。我们发现,MN的大多数性能挑战来自将内存多维数据集绑定在一起的互连网络。特别是,用于路由通过MNs的仲裁方案、NVM与DRAM的比例以及所使用的特定拓扑对性能和能源结果有巨大影响。我们的初步分析表明,在MN中引入非易失性内存在内存阵列延迟和网络延迟之间提供了一个独特的权衡。我们观察到,在MN中以特定顺序放置NVM多维数据集可以通过将网络大小/直径减小到一定的NVM与DRAM之比来提高性能。新颖的MN拓扑和仲裁方案还通过减少MN中请求和响应的跳数来提供性能和能量增量。根据我们的分析,我们介绍了三种技术来解决MN延迟问题:(1)基于距离的仲裁方案,以改善整个网络的排队延迟;(2)源自经典数据结构的跳表拓扑,以改善网络延迟和链路使用;(3)MetaCube,一种更密集的内存立方体,利用先进的封装技术通过减少MN大小来改善延迟。
High-performance computing, enterprise, and datacenter servers are driving demands for higher total memory capacity as well as memory performance. Memory “cubes” with high per-package capacity (from 3D integration) along with high-speed point-to-point interconnects provide a scalable memory system architecture with the potential to deliver both capacity and performance. Multiple such cubes connected together can form a “Memory Network” (MN), but the design space for such MNs is quite vast, including multiple topology types and multiple memory technologies per memory cube. In this work, we first analyze several MN topologies with different mixes of memory package technologies to understand the key tradeoffs and bottlenecks for such systems. We find that most of a MN's performance challenges arise from the interconnection network that binds the memory cubes together. In particular, arbitration schemes used to route through MNs, ratio of NVM to DRAM, and specific topologies used have dramatic impact on performance and energy results. Our initial analysis indicates that introducing non-volatile memory to the MN presents a unique tradeoff between memory array latency and network latency. We observe that placing NVM cubes in a specific order in the MN improves performance by reducing the network size/diameter up to a certain NVM to DRAM ratio. Novel MN topologies and arbitration schemes also provide performance and energy deltas by reducing the hop count of requests and response in the MN. Based on our analyses, we introduce three techniques to address MN latency issues: (1) Distance-based arbitration scheme to improve queuing latencies throughout the network, (2) skip-list topology, derived from the classic data structure, to improve network latency and link usage, and (3) the MetaCube, a denser memory cube that leverages advanced packaging technologies to improve latency by reducing MN size.