Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion

Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion
复制标题

DOI:
10.1145/3582016.3582032
复制
发表时间:
2023-03
期刊:
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3
影响因子:
--
通讯作者:
Zhengrong Wang;Christopher Liu;Aman Arora;L. John;Tony Nowatzki
Zhengrong Wang;Christopher Liu;Aman Arora;L. John;Tony Nowatzki
中科院分区:
其他
文献类型:
--
作者:
Zhengrong Wang;Christopher Liu;Aman Arora;L. John;Tony Nowatzki

文献摘要

相似文献

具有大型末级缓存的内存中计算有望显着缓解数据移动瓶颈并带来大量位线级并行化机会。然而,其独特的执行模型的关键挑战仍然没有解决:自动并行化,透明地编排位串行逻辑的数据转换/对齐/广播,以及混合内存/近内存计算。最重要的是,解决方案应该是程序员友好的,跨平台的可移植性。我们的关键创新是一个执行模型和中间表示(IR),支持混合CPU核心,内存和近内存处理。我们的IR是张量图(tDFG),它是内存和近内存计算的统一表示。tDFG公开张量数据结构信息,以便硬件和运行时可以自动编排用于位串行执行的数据管理,包括运行时数据布局转换。为了实现微架构的可移植性,我们使用了两阶段,基于JIT的编译方法来动态降低tDFG到内存中的命令。我们的设计,无限流,周期精确的模拟器上进行评估。在使用fp 32的数据处理工作负载中,与最先进的近内存计算技术相比,它实现了2.6倍的加速和75%的流量减少,并具有2.4倍的能效。
In-memory computing with large last-level caches is promising to dramatically alleviate data movement bottlenecks and expose massive bitline-level parallelization opportunities. However, key challenges from its unique execution model remain unsolved: automated parallelization, transparently orchestrating data transposition/alignment/broadcast for bit-serial logic, and mixing in-/near-memory computing. Most importantly, the solution should be programmer friendly and portable across platforms. Our key innovation is an execution model and intermediate representation (IR) that enables hybrid CPU-core, in-memory, and near-memory processing. Our IR is the tensor dataflow graph (tDFG), which is a unified representation of in-memory and near-memory computation. The tDFG exposes tensor-data structure information so that the hardware and runtime can automatically orchestrate data management for bitserial execution, including runtime data layout transformations. To enable microarchitecture portability, we use a two-phase, JIT-based compilation approach to dynamically lower the the tDFG to in-memory commands. Our design, infinity stream, is evaluated on a cycle-accurate simulator. Across data-processing workloads with fp32, it achieves 2.6× speedup and 75% traffic reduction over a state-of-the-art near-memory computing technique, with 2.4× energy efficiency.