ValueExpert: exploring value patterns in GPU-accelerated applications

ValueExpert: exploring value patterns in GPU-accelerated applications
复制标题

DOI:
10.1145/3503222.3507708
复制
发表时间:
2022-02
期刊:
Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
K. Zhou;Yueming Hao;J. Mellor-Crummey;Xiaozhu Meng;Xu Liu
K. Zhou;Yueming Hao;J. Mellor-Crummey;Xiaozhu Meng;Xu Liu
中科院分区:
其他
文献类型:
--
作者:
K. Zhou;Yueming Hao;J. Mellor-Crummey;Xiaozhu Meng;Xu Liu

文献摘要

相似文献

通用GPU已在现代计算系统中很普遍,以加速许多领域的应用,包括机器学习,高性能计算和自动驾驶。但是,GPU加速应用中效率低下,这阻止了它们获得裸机的性能。性能工具在理解复杂代码库中的性能低效率方面起着重要作用。许多GPU性能工具都会查明耗时的代码,并提供高级的性能见解,但忽略了一个重要的性能问题 - 与价值相关的效率低下,这在许多GPU代码库中都存在。在本文中,我们提出了ValueAppert,这是一种新型工具,可以查明GPU应用中与价值相关的效率低下。 ValueExpert监视应用程序执行,以捕获GPU内核中每个负载和存储操作产生和使用的值,识别多个值模式,并提供直观的优化指南。我们解决了从许多GPU线程收集,维护和分析大量性能数据时的系统性挑战,以使价值Expert适用于复杂的应用程序。我们评估了广泛调整良好的基准和应用的价值专家,包括Pytorch,Darknet,Lammps,Castro等。 ValueExpert能够识别以前未知的绩效问题,并为非平凡性能改进提供建议,通常少于五行代码更改。我们通过应用程序开发人员和上游修复库来验证我们的优化。
General-purpose GPUs have become common in modern computing systems to accelerate applications in many domains, including machine learning, high-performance computing, and autonomous driving. However, inefficiencies abound in GPU-accelerated applications, which prevent them from obtaining bare-metal performance. Performance tools play an important role in understanding performance inefficiencies in complex code bases. Many GPU performance tools pinpoint time-consuming code and provide high-level performance insights but overlook one important performance issue---value-related inefficiencies, which exist in many GPU code bases. In this paper, we present ValueExpert, a novel tool to pinpoint value-related inefficiencies in GPU applications. ValueExpert monitors application execution to capture values produced and used by each load and store operation in GPU kernels, recognizes multiple value patterns, and provides intuitive optimization guidance. We address systemic challenges in collecting, maintaining, and analyzing voluminous performance data from many GPU threads to make ValueExpert applicable to complex applications. We evaluate ValueExpert on a wide range of well-tuned benchmarks and applications, including PyTorch, Darknet, LAMMPS, Castro, and many others. ValueExpert is able to identify previously unknown performance issues and provide suggestions for nontrivial performance improvements with typically less than five lines of code changes. We verify our optimizations with application developers and upstream fixes to their repositories.