Accelerating Applications using Edge Tensor Processing Units

Accelerating Applications using Edge Tensor Processing Units
复制标题

DOI:
10.1145/3458817.3476177
复制
发表时间:
2021-06
期刊:
SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Kuan-Chieh Hsu;Hung-Wei Tseng
Kuan-Chieh Hsu;Hung-Wei Tseng
中科院分区:
其他
文献类型:
--
作者:
Kuan-Chieh Hsu;Hung-Wei Tseng

文献摘要

相似文献

神经网络(NN)加速器已集成到广泛的计算机系统中,以适应人工智能(AI)和机器学习(ML)应用的快速增长的需求。 NN加速器分享了为多维张量数据运行提供本地硬件支持的想法。因此,NN加速器是理论上张量处理器,可以改善使用张量作为输入/输出的任何问题的系统性能。不幸的是,市售的NN加速器仅通过AI/ML特异性接口揭示计算功能。此外,NN加速器揭示了很少的硬件设计细节,因此应用程序无法轻易利用Tensor操作NN加速器提供的。本文介绍了张量处理单元(GPTPU)的通用计算,这是一个开源的开放式结构框架,使开发人员和研究社区可以发现NN加速器为应用程序启用的机会。 GPTPU包括一个功能强大的编程接口,具有有效的运行时系统级别的支持类似于GPGPU计算中CUDA/OPENCL的编程界面,以弥合应用程序需求之间的差距,并且不匹配的硬件/软件接口之间的差距。我们构建了GPTPU机器,使用Edge Tensor处理单元(Edge TPU),该单元广泛可用,并代表了许多商业NN加速器。我们确定了几种新颖的用例,并重新审视了算法。通过利用底层边缘TPU执行基于张量的计算机内核,我们的结果表明,GPTPU可以在高端CPU上实现2.46倍的速度,并将能源消耗降低40%。
Neural network (NN) accelerators have been integrated into a wide-spectrum of computer systems to accommodate the rapidly growing demands for artificial intelligence (AI) and machine learning (ML) applications. NN accelerators share the idea of providing native hardware support for operations on multidimensional tensor data. Therefore, NN accelerators are theoretically tensor processors that can improve system performance for any problem that uses tensors as inputs/outputs. Unfortunately, commercially available NN accelerators only expose computation capabilities through AI/ML-specific interfaces. Furthermore, NN accelerators reveal very few hardware design details, so applications cannot easily leverage the tensor operations NN accelerators provide. This paper introduces General-Purpose Computing on Tensor Processing Units (GPTPU), an open-source, open-architecture framework that allows the developer and research communities to discover opportunities that NN accelerators enable for applications. GPTPU includes a powerful programming interface with efficient runtime system-level support-similar to that of CUDA/OpenCL in GPGPU computing-to bridge the gap between application demands and mismatched hardware/software interfaces. We built GPTPU machine uses Edge Tensor Processing Units (Edge TPUs), which are widely available and representative of many commercial NN accelerators. We identified several novel use cases and revisited the algorithms. By leveraging the underlying Edge TPUs to perform tensor-algorithm-based compute kernels, our results reveal that GPTPU can achieve a 2.46x speedup over high-end CPUs and reduce energy consumption by 40%.