A Tensor Processing Framework for CPU-Manycore Heterogeneous Systems

A Tensor Processing Framework for CPU-Manycore Heterogeneous Systems
复制标题

DOI:
10.1109/tcad.2021.3103825
复制
发表时间:
2022-06
影响因子:
2.9
通讯作者:
Lin Cheng;Peitian Pan;Zhongyuan Zhao;Krithik Ranjan;Jack Weber;Bandhav Veluri;Seyed Borna Ehsani;Max Ruttenberg;Dai Cheol Jung;Preslav Ivanov;D. Richmond;M. Taylor;Zhiru Zhang;C. Batten
Lin Cheng;Peitian Pan;Zhongyuan Zhao;Krithik Ranjan;Jack Weber;Bandhav Veluri;Seyed Borna Ehsani;Max Ruttenberg;Dai Cheol Jung;Preslav Ivanov;D. Richmond;M. Taylor;Zhiru Zhang;C. Batten
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lin Cheng;Peitian Pan;Zhongyuan Zhao;Krithik Ranjan;Jack Weber;Bandhav Veluri;Seyed Borna Ehsani;Max Ruttenberg;Dai Cheol Jung;Preslav Ivanov;D. Richmond;M. Taylor;Zhiru Zhang;C. Batten

文献摘要

被引文献

相似文献

未来的CPU - 多核异质系统可以通过将数千个简单,独立,节能的核心整合到单个模具中来提供高峰值吞吐量。 :1)许多核心处理器依靠简单的硬件对软件程序员提出了重大需求,2)许多核心处理器使用努力容忍的固定核心长期记忆潜伏期,以解决许多核心的可编程性挑战,基于pytorch提供了一个稀疏的张量处理框架,使域专家能够轻松地加速CPU-ManyCore Ockore Onshelf Workloads。挑战,我们使用扩展的Pytorch框架来探索脱钩/执行(DAE)软件和硬件机制的潜力我们使用Pytorch运算符微型计算和现实世界的Pytorch工作负载来评估我们的技术,在详细的寄存器 - 转移级模型上运行了128核Many Core Architecture。稀疏张量的工作负载表明,与18核较高的CPU基线相比,这些工作量可以实现大约2– $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6 \ $ 6与通用图形处理单元相比,标准化的吞吐量和提高的能源效率。
Future CPU-manycore heterogeneous systems can provide high peak throughput by integrating thousands of simple, independent, energy-efficient cores in a single die. However, there are two key challenges to translating this high peak throughput into improved end-to-end workload performance: 1) manycore co-processors rely on simple hardware putting significant demands on the software programmer and 2) manycore co-processors use in-order cores that struggle to tolerate long memory latencies. To address the manycore programmability challenge, this article presents a dense and sparse tensor processing framework based on PyTorch that enables domain experts to easily accelerate off-the-shelf workloads on CPU-manycore heterogeneous systems. To address the manycore memory latency challenge, we use our extended PyTorch framework to explore the potential for decoupled access/execute (DAE) software and hardware mechanisms. More specifically, we propose two software-only techniques, naïve-software DAE and systolic-software DAE, along with a lightweight hardware access accelerator to further improve area-normalized throughput. We evaluate our techniques using a combination of PyTorch operator microbenchmarking and real-world PyTorch workloads running on a detailed register-transfer-level model of a 128-core manycore architecture. Our evaluation on three real-world dense and sparse tensor workloads suggests these workloads can achieve approximately 2– $6\times $ performance improvement when scaled to a future 2000-core CPU-manycore heterogeneous system compared to an 18-core out-of-order CPU baseline, while potentially achieving higher area-normalized throughput and improved energy efficiency compared to general-purpose graphics processing units.