In-Memory Data Parallel Processor

In-Memory Data Parallel Processor
复制标题

DOI:
10.1145/3173162.3173171
复制
发表时间:
2018-03
期刊:
Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Daichi Fujiki;S. Mahlke;R. Das
Daichi Fujiki;S. Mahlke;R. Das
中科院分区:
其他
文献类型:
--
作者:
Daichi Fujiki;S. Mahlke;R. Das

文献摘要

被引文献

相似文献

非易失性存储器(NVM)的最新发展为内存计算开辟了新的视野。尽管计算NVM提供了显着的性能增益,但以前的工作依赖于将专用内核手动映射到内存阵列,使得执行更一般的工作负载变得不可行。我们提出了一个可编程的内存处理器架构和数据并行编程框架来解决这个问题。内存中处理器的效率来自两个来源:大规模并行和减少数据移动。紧凑指令集为存储器阵列提供通用计算能力。建议的编程框架旨在通过合并数据流和向量处理的概念来利用硬件中的底层并行性。为了方便内存编程,我们开发了一个编译框架,该框架接受TensorFlow输入并为我们的内存处理器生成代码。我们的结果表明,对于Parsec的一组应用程序,多核CPU服务器的加速比为7.5倍,对于Rodinia的一组基准测试,服务器级GPU的加速比为763倍。
Recent developments in Non-Volatile Memories (NVMs) have opened up a new horizon for in-memory computing. Despite the significant performance gain offered by computational NVMs, previous works have relied on manual mapping of specialized kernels to the memory arrays, making it infeasible to execute more general workloads. We combat this problem by proposing a programmable in-memory processor architecture and data-parallel programming framework. The efficiency of the proposed in-memory processor comes from two sources: massive parallelism and reduction in data movement. A compact instruction set provides generalized computation capabilities for the memory array. The proposed programming framework seeks to leverage the underlying parallelism in the hardware by merging the concepts of data-flow and vector processing. To facilitate in-memory programming, we develop a compilation framework that takes a TensorFlow input and generates code for our in-memory processor. Our results demonstrate 7.5x speedup over a multi-core CPU server for a set of applications from Parsec and 763x speedup over a server-class GPU for a set of Rodinia benchmarks.