CASH: compiler assisted hardware design for improving DRAM energy efficiency in CNN inference

CASH: compiler assisted hardware design for improving DRAM energy efficiency in CNN inference
复制标题

DOI:
10.1145/3357526.3357536
复制
发表时间:
2019-09
期刊:
Proceedings of the International Symposium on Memory Systems
影响因子:
--
通讯作者:
Anup Sarma;Huaipan Jiang;Ashutosh Pattnaik;Jagadish B. Kotra;M. Kandemir;C. Das
Anup Sarma;Huaipan Jiang;Ashutosh Pattnaik;Jagadish B. Kotra;M. Kandemir;C. Das
中科院分区:
其他
文献类型:
--
作者:
Anup Sarma;Huaipan Jiang;Ashutosh Pattnaik;Jagadish B. Kotra;M. Kandemir;C. Das

文献摘要

相似文献

机器学习(ML)和深度学习应用的出现导致了大量硬件加速器和并行架构的架构优化技术的发展。这在一定程度上是由于ML工作负载所表现出的规律性和并行性,特别是卷积神经网络(CNN)。然而,CPU仍然是当今数据中心的主要计算结构之一,因此也被广泛部署用于推理任务。随着CNN变得越来越大,基于CPU的系统的固有局限性变得越来越明显,特别是在主存储器数据移动方面。在本文中,我们提出了CASH,编译器辅助的硬件解决方案,消除了冗余的数据移动和从主存储器,因此,减少了主存储器的带宽和能耗。我们对一组四种不同的最先进的CNN工作负载进行的实验评估表明,CASH平均分别提供了约40%和约18%的主存带宽和能耗降低。
The advent of machine learning (ML) and deep learning applications has led to the development of a multitude of hardware accelerators and architectural optimization techniques for parallel architectures. This is due in part to the regularity and parallelism exhibited by the ML workloads, especially convolutional neural networks (CNNs). However, CPUs continue to be one of the dominant compute fabric in data-centers today, thereby also being widely deployed for inference tasks. As CNNs grow larger, the inherent limitations of a CPU-based system become apparent, specifically in terms of main memory data movement. In this paper, we present CASH, a compiler-assisted hardware solution that eliminates redundant data-movement to and from the main memory and, therefore, reduces main memory bandwidth and energy consumption. Our experimental evaluations on a set of four different state-of-the-art CNN workloads indicate that CASH provides, on average, ~40% and ~18% reductions in main memory bandwidth and energy consumption, respectively.