CORUSCANT: Fast Efficient Processing-in-Racetrack Memories

CORUSCANT: Fast Efficient Processing-in-Racetrack Memories
复制标题

DOI:
10.1109/micro56248.2022.00060
复制
发表时间:
2022-10
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
S. Ollivier;Stephen Longofono;Prayash Dutta;J. Hu;S. Bhanja;A. Jones
S. Ollivier;Stephen Longofono;Prayash Dutta;J. Hu;S. Bhanja;A. Jones
中科院分区:
其他
文献类型:
--
作者:
S. Ollivier;Stephen Longofono;Prayash Dutta;J. Hu;S. Bhanja;A. Jones

文献摘要

被引文献

相似文献

现代应用程序对数据需求的增长给现代系统带来了重大挑战,导致了“内存墙”问题。自旋电子畴壁存储器(DWM)具有接近静态随机存取存储器(SRAM)的读/写性能、节能且具有非易失性、有极高存储密度的潜力,并且没有明显的耐久性限制。然而,DWM的优势无法直接解决数据访问延迟以及内存总线带宽的吞吐量限制问题。内存内处理(PIM)是一种流行的解决方案,它通过将计算直接转移到内存来减少内存与处理器之间通信的需求。PIM已在多种技术中被提出,包括动态随机存取存储器(DRAM)、相变存储器(PCM)、电阻式存储器(ReRAM)和自旋转移矩存储器(STT - MRAM)。DRAM PIM为一组受限的双操作数按位批量操作提供了解决方案。PCM和ReRAM中的PIM引发了人们对其有效耐久性的担忧,而STT - MRAM中的PIM对于主存应用来说密度不足。我们提出了CORUSCANT,一种基于DWM的内存内计算解决方案,它利用了DWM纳米线的特性,并使其能够用作多态门。通常情况下,DWM是通过在访问点施加与纳米线正交的自旋极化电流来访问单个位的,而沿着DWM纳米线的横向访问可以区分纳米线中多个位的总电阻,类似于多级单元。CORUSCANT利用这种横向读取直接提供多操作数按位批量逻辑。利用横向访问所实现的这种多操作数概念,CORUSCANT提供的技术在进行多操作数加法和双操作数乘法时比先前的数字PIM解决方案高效得多。对于利用按位批量操作的查询应用,与领先的DRAM PIM技术相比,CORUSCANT的速度提高了1.6倍。与用于DWM的领先PIM技术相比,对于8位加法和乘法,CORUSCANT的性能分别提高了6.9倍和2.3倍,能耗分别降低了5.5倍和3.4倍。对于算术密集型基准测试,与非PIM的DWM相比,CORUSCANT在面积开销为10%的情况下,访问延迟降低了2.1倍,能耗降低了25.2倍。
The growth in data needs of modern applications has created significant challenges for modern systems leading to a “memory wall.” Spintronic Domain-Wall Memory (DWM), provides near-SRAM read/write performance, energy savings and non-volatility, potential for extremely high storage density, and does not have significant endurance limitations. However, DWM’s benefits cannot directly address data access latency and throughput limitations of memory bus bandwidth. Processing-inmemory (PIM) is a popular solution to reduce the demands of memory-to-processor communication by offloading computation directly to the memory. PIM has been proposed in multiple technologies including DRAM, Phase-change memory (PCM), resistive memory (ReRAM), and Spin-Transfer Torque Memory (STT-MRAM). DRAM PIM provides solutions for a restricted set of two operand bulk-bitwise operations. PIM in PCM and ReRAM raise concerns about their effective endurance and PIM in STT-MRAM has insufficient density for main-memory applications. We propose CORUSCANT, a DWM-based in-memory computing solution that leverages the properties of DWM nanowires and allows them to serve as polymorphic gates. While normally DWM is accessed by applying spin polarized currents orthogonal to the nanowire at access points to read individual bits, transverse access along the DWM nanowire allows the differentiation of the aggregate resistance of multiple bits in the nanowire, akin to a multi-level cell. CORUSCANT leverages this transverse reading to directly provide multi-operand bulk-bitwise logic. Leveraging this multi-operand concept enabled by transverse access, CORUSCANT provides techniques to conduct multi-operand addition and two operand multiplication much more efficiently than prior digital PIM solutions. CORUSCANT provides a 1.6 × speedup compared to the leading DRAM PIM technique for query applications that leverage bulk bitwise operations. Compared to the leading PIM technique for DWM, CORUSCANT improves performance by 6.9 ×, 2.3 × and energy by 5.5 ×, 3.4 × for 8-bit addition and multiplication, respectively. For arithmetic heavy benchmarks, CORUSCANT reduces access latency by 2.1 ×, while decreasing energy consumption by 25.2 × for a 10% area overhead versus non-PIM DWM.