Thread-Aware Cache Simulator for HPC Application Tuning

Thread-Aware Cache Simulator for HPC Application Tuning
复制标题

用于 HPC 应用程序调优的线程感知缓存模拟器

DOI:
10.1109/candarw53999.2021.00078
复制
发表时间:
2021
期刊:
In Proceedings of 12th International Workshop on Advances in Networking and Computing (WANC 2021)
影响因子:
--
通讯作者:
Sato Yukinori
Sato Yukinori
中科院分区:
--
文献类型:
--
作者:
Chugo Kazuki;Sato Yukinori

文献摘要

相似文献

在本文中,我们提出了一个线程感知和并行缓存模拟器,以协助性能调优工作流程的HPC应用程序。我们的模拟器模仿富士通A64FX处理器的多级缓存的行为,并能够检测缓存的冲突未命中。由于已知通过调整技术(例如阵列填充)来减少高速缓存冲突未命中,因此冲突未命中的检测对于调整HPC应用非常有用。为了加快多线程代码的模拟,我们研究了使用动态二进制插装框架实现的在线缓存模拟器的设计空间。为了对单个CPU内的多个L1数据缓存进行建模并将它们映射到线程,我们构建了两种设计:1)单片模拟器,其中通过观察到的存储器轨迹的消息传递通信,使用分离的进程独立地执行CPU的所有L1数据高速缓存模拟,2)并行模拟器,其中L1数据高速缓存模拟在运行原始应用程序的同一线程中作为插装代码来执行。我们评估的准确性和运行时间,我们的模拟器使用多线程的基准程序与A64FX处理器的机器上。通过比较内存引用和L1数据缓存未命中与富士通分析器的数量,我们确认,我们的模拟器可以准确地模拟底层缓存的行为。评估结果表明,并行模拟器设计具有可扩展性,根据分配给每个应用程序的线程数,其执行时间总是短于单片模拟器设计。这些结果表明,并行模拟器的设计是有前途的,以协助多线程HPC应用程序的性能调优工作流程。
In this paper, we present a thread-aware and parallel cache simulator to assist a performance tuning workflow of HPC applications. Our simulator mimics the behavior of multilevel cache of Fujitsu A64FX processor and enables to detect conflict misses of caches. Since cache conflict misses are known to be reduced by tuning techniques such as array padding, the detection of conflict misses is very useful for tuning HPC applications. To speed up simulations for multi-threaded code, we investigate a design space of the online cache simulator implemented using a dynamic binary instrumentation framework. To model multiple L1 data caches within a single CPU and map them to threads, we build two designs: 1) monolithic simulator, where all of L1 data cache simulations for a CPU are independently performed using a separated process by message passing communication of observed memory traces, 2) parallel simulator, where an L1 data cache simulation is performed in the same thread running the original application as instrumentation code. We evaluate the accuracy and the runtime of our simulator using multi-threaded benchmark programs on the machine with an A64FX processor. By comparing the number of memory references and L1 data cache misses with Fujitsu profiler, we confirm that our simulator can accurately model the underlying cache behavior. The evaluation results show that the parallel simulator design has scalability according to the number of threads assigned for each application, and its execution time is always shorter than that of the monolithic simulator design. These results indicate that the parallel simulator design is promising for assisting a performance tuning workflow of multi-threaded HPC applications.