Accelerating Low Bit-width Neural Networks at the Edge, PIM or FPGA: A Comparative Study

Accelerating Low Bit-width Neural Networks at the Edge, PIM or FPGA: A Comparative Study
复制标题

DOI:
10.1145/3583781.3590213
复制
发表时间:
2023-06
期刊:
Proceedings of the Great Lakes Symposium on VLSI 2023
影响因子:
--
通讯作者:
Nakul Kochar;Lucas Ekiert;Deniz Najafi;Deliang Fan;Shaahin Angizi
Nakul Kochar;Lucas Ekiert;Deniz Najafi;Deliang Fan;Shaahin Angizi
中科院分区:
其他
文献类型:
--
作者:
Nakul Kochar;Lucas Ekiert;Deniz Najafi;Deliang Fan;Shaahin Angizi

文献摘要

相似文献

深度神经网络(DNN)加速与边缘的数字内存处理(PIM)平台是一个积极探索的领域,不仅有很大的潜力解决内存墙瓶颈,而且与冯·诺依曼架构相比,提供了性能改进的订单。另一方面,基于FPGA的边缘计算已被视为加速计算密集型工作负载的潜在解决方案。在这项工作中,采用低位宽的神经网络,我们进行了坚实的和比较的推理性能分析最近的处理在SRAM流片与低资源的FPGA板和高性能的GPU,为研究界提供了一个指南。我们探索并强调这些边缘候选人的关键架构约束,影响其整体性能。我们的实验数据表明,与CIFAR-10数据集上的待测FPGA相比,SRAM中的处理可以获得高达160倍的速度提升和高达228倍的效率(img/s/W)。
Deep Neural Network (DNN) acceleration with digital Processing-in-Memory (PIM) platforms at the edge is an actively-explored domain with great potential to not only address memory-wall bottlenecks but to offer orders of performance improvement in comparison to the von-Neumann architecture. On the other side, FPGA-based edge computing has been followed as a potential solution to accelerate compute-intensive workloads. In this work, adopting low-bit-width neural networks, we perform a solid and comparative inference performance analysis of a recent processing-in-SRAM tape-out with a low-resource FPGA board and a high-performance GPU to provide a guideline for the research community. We explore and highlight the key architectural constraints of these edge candidates that impact their overall performance. Our experimental data demonstrate that the processing-in-SRAM can obtain up to ~160x speed-up and up to 228x higher efficiency (img/s/W) compared to the under-test FPGA on the CIFAR-10 dataset.