TPUPoint: Automatic Characterization of Hardware-Accelerated Machine-Learning Behavior for Cloud Computing

TPUPoint: Automatic Characterization of Hardware-Accelerated Machine-Learning Behavior for Cloud Computing
复制标题

DOI:
10.1109/ispass51385.2021.00048
复制
发表时间:
2021-03
期刊:
2021 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
影响因子:
--
通讯作者:
Abenezer Wudenhe;Hung-Wei Tseng
Abenezer Wudenhe;Hung-Wei Tseng
中科院分区:
其他
文献类型:
--
作者:
Abenezer Wudenhe;Hung-Wei Tseng

文献摘要

相似文献

随着机器学习(ML)在数据中心中的工作量的份额迅速增加,云提供商开始合并加速器,例如张量处理单元(TPU),以提高应用程序的能源效率。但是,在不优化应用程序参数的情况下,用户可能不足以利用加速器并最终浪费能量和金钱。本文介绍了TPUPOINT,以促进基于TPU的云平台上有效应用的开发。 TPUPOINT会自动将重复模式分为阶段,并确定每个阶段中最关键的关键操作。此外,TPUPOINT可以将阶段与检查点相关联,以允许在应用程序中进行快速发展,从而大大减少了用于优化应用程序的时间和金钱。通过在各种代表性的ML工作负载上运行TPUPOINT,我们发现计算不再是最耗时的操作。取而代之的是,交换和重新调整数据的进化和重塑操作变得最重要。 TPupoints优势显着增加了发现最佳参数的潜力,以快速平衡将数据的复杂工作负载管线平衡到系统中,重新格式化数据和计算结果。
With the share of machine learning (ML) workloads in data centers rapidly increasing, cloud providers are beginning to incorporate accelerators such as tensor processing units (TPUs) to improve the energy-efficiency of applications. However, without optimizing application parameters, users may underutilize accelerators and end up wasting energy and money. This paper presents TPUPoint to facilitate the development of efficient applications on TPU-based cloud platforms. TPUPoint automatically classifies repetitive patterns into phases and identifies the most timing-critical operations in each phase. Further, TPUPoint can associate phases with checkpoints to allow fast-forwarding in applications, thereby significantly reducing the time and money spent optimizing applications. By running TPUPoint on a wide array of representative ML workloads, we found that computation is no longer the most time-consuming operation; instead, the infeed and reshape operations, which exchange and realign data, become most significant. TPUPoints advantages significantly increase the potential for discovering optimal parameters to quickly balance the complex workload pipeline of feeding data into a system, reformatting the data, and computing results.