NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip Accelerators

NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip Accelerators
复制标题

DOI:
10.1109/isca52012.2021.00083
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Zhanhong Tan;Hongyu Cai;Runpei Dong;Kaisheng Ma
Zhanhong Tan;Hongyu Cai;Runpei Dong;Kaisheng Ma
中科院分区:
其他
文献类型:
--
作者:
Zhanhong Tan;Hongyu Cai;Runpei Dong;Kaisheng Ma

文献摘要

被引文献

相似文献

机器学习的革命对计算资源提出了前所未有的需求,促使单个单片芯片上安装更多晶体管,这在后摩尔时代是不可持续的。具有小型功能芯片(称为小芯片)的多芯片集成可以降低制造成本,提高制造良率,并实现不同系统规模的芯片级重用。此类多芯片系统上的 DNN 工作负载映射和硬件设计空间探索至关重要,但在当前阶段还缺少。这项工作提供了一个分层分析框架来描述多芯片加速器上的 DNN 映射并分析通信开销。基于这个框架,我们提出了一种名为 NN-Baton 的自动化工具,具有预设计流程和后设计流程。预设计流程旨在指导针对目标工作负载的给定面积和性能预算的小芯片粒度探索。后设计流程重点关注层次结构中不同计算级别(封装、小芯片和核心)的工作负载编排。与 Simba 相比,NN-Baton 生成的映射策略在相同计算和内存配置下可节省 22.5%∼44% 的能量。架构探索表明,面积是 Chiplet 粒度的决定性因素。对于 2 mm2 小芯片面积限制下的 2048-MAC 系统,具有 4 个内核和 16 个通道的 8 尺寸矢量 MAC 的 4 小芯片实现始终是多个基准测试中的首选计算分配。相反,层次结构中的最佳内存分配策略通常取决于神经网络模型。
The revolution of machine learning poses an unprecedented demand for computation resources, urging more transistors on a single monolithic chip, which is not sustainable in the Post-Moore era. The multichip integration with small functional dies, called chiplets, can reduce the manufacturing cost, improve the fabrication yield, and achieve die-level reuse for different system scales. DNN workload mapping and hardware design space exploration on such multichip systems are critical, but missing in the current stage.This work provides a hierarchical and analytical framework to describe the DNN mapping on a multichip accelerator and analyze the communication overhead. Based on this framework, we propose an automatic tool called NN-Baton with a pre-design flow and a post-design flow. The pre-design flow aims to guide the chiplet granularity exploration with given area and performance budgets for the target workload. The post-design flow focuses on the workload orchestration on different computation levels -package, chiplet, and core - in the hierarchy. Compared to Simba, NN-Baton generates mapping strategies that save 22.5%∼44% energy under the same computation and memory configurations.The architecture exploration demonstrates that area is a decisive factor for the chiplet granularity. For a 2048-MAC system under a 2 mm2 chiplet area constraint, the 4-chiplet implementation with 4 cores and 16 lanes of 8-size vector-MAC is always the top-pick computation allocation across several benchmarks. In contrast, the optimal memory allocation policy in the hierarchy typically depends on the neural network models.