Merge Network for a Non-Von Neumann Accumulate Accelerator in a 3D Chip

Merge Network for a Non-Von Neumann Accumulate Accelerator in a 3D Chip
复制标题

3D 芯片中非冯诺依曼累积加速器的合并网络

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Rebooting Computing
影响因子:
--
通讯作者:
T. Krishna
T. Krishna
中科院分区:
--
文献类型:
--
作者:
Anirudh Jain;S. Srikanth;E. Debenedictis;T. Krishna

文献摘要

被引文献

相似文献

逻辑内存集成有助于缓解冯·诺依曼瓶颈,这使得一类新的架构能够帮助加速稀疏数据流上的图形分析和操作。它们利用合并网络作为关键的计算单元。这样的网络是高度并行的,并且当使用双调算法时,它们的性能随着逻辑和存储器之间的更紧密耦合而增加。本文提出了一种高能效的片上网络架构,用于合并使用字并行和位串行范式的键值对。所提出的架构能够合并两行高带宽存储器(HBM)的数据,其方式与从这样的行的阅读和写回这样的行完全重叠。此外,当与基于朴素交叉开关的设计相比时,它们的能量消耗大约低一个数量级。
Logic-memory integration helps mitigate the von Neumann bottleneck, and this has enabled a new class of architectures that helps accelerate graph analytics and operations on sparse data streams. These utilize merge networks as a key unit of computation. Such networks are highly parallel and their performance increases with tighter coupling between logic and memory when a bitonic algorithm is used. This paper presents energy-efficient on-chip network architectures for merging key-value pairs using both word-parallel and bit-serial paradigms. The proposed architectures are capable of merging two rows of high bandwidth memory (HBM)worth of data in a manner that is completely overlapped with the reading from and writing back to such a row. Furthermore, their energy consumption is about an order of magnitude lower when compared to a naive crossbar based design.