Hatchet: pruning the overgrowth in parallel profiles

Hatchet: pruning the overgrowth in parallel profiles
复制标题

斧头:修剪平行轮廓的过度生长

DOI:
10.1145/3295500.3356219
复制
发表时间:
2019
期刊:
Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
T. Gamblin
T. Gamblin
中科院分区:
--
文献类型:
--
作者:
A. Bhatele;S. Brink;T. Gamblin

文献摘要

被引文献

相似文献

性能分析对于消除并行代码中的可伸缩性瓶颈至关重要。有许多分析工具可以检测代码并收集性能数据。然而,通用的、易于使用的、可编程的分析和可视化工具是有限的。在本文中,我们关注结构化分析数据的分析,例如通过调用上下文树或代码中嵌套的区域计时器获得的数据。我们提供了一组基于pandas数据分析库的技术和操作,以支持并行配置文件的分析。我们在一个名为Hatchet的基于python的库中实现了这些技术,该库允许对结构化数据进行过滤、聚合和修剪。使用从分析并行代码中获得的性能数据集,我们演示了使用几行Hatchet代码可重复地执行常见的性能分析任务。Hatchet将现代数据科学工具的强大功能用于性能分析。
Performance analysis is critical for eliminating scalability bottlenecks in parallel codes. There are many profiling tools that can instrument codes and gather performance data. However, analytics and visualization tools that are general, easy to use, and programmable are limited. In this paper, we focus on the analytics of structured profiling data, such as that obtained from calling context trees or nested region timers in code. We present a set of techniques and operations that build on the pandas data analysis library to enable analysis of parallel profiles. We have implemented these techniques in a Python-based library called Hatchet that allows structured data to be filtered, aggregated, and pruned. Using performance datasets obtained from profiling parallel codes, we demonstrate performing common performance analysis tasks reproducibly with a few lines of Hatchet code. Hatchet brings the power of modern data science tools to bear on performance analysis.