CachePerf: A Unified Cache Miss Classifier via Hybrid Hardware Sampling

CachePerf: A Unified Cache Miss Classifier via Hybrid Hardware Sampling
复制标题

CachePerf:通过混合硬件采样的统一缓存未命中分类器

DOI:
10.1145/3547353.3526954
复制
发表时间:
2022
期刊:
ACM SIGMETRICS Performance Evaluation Review
影响因子:
--
通讯作者:
Liu, Tongping
Liu, Tongping
中科院分区:
--
文献类型:
--
作者:
Zhou, Jin;Tang, Steven;Yang, Hanmei;Liu, Tongping

文献摘要

参考文献

相似文献

缓存在决定应用程序的性能方面起着关键作用,无论是同构还是异构架构上的顺序程序还是并发程序。修复缓存缺失需要了解缓存缺失的来源和类型。然而,即使经过几十年的研究,这仍然是一个未解决的问题。本文提出了一个统一的分析工具——CachePerf——它可以正确地识别不同类型的缓存缺失,区分分配器引起的问题和应用程序的问题,并在不太影响性能的情况下排除次要问题。CachePerf背后的核心思想是一种混合采样方案:它采用基于pmu的粗粒度采样来选择很少的易受影响的指令(经常缓存丢失),然后采用基于断点的细粒度采样来收集这些指令的内存访问模式。根据我们的评估,CachePerf在正确识别缓存丢失类型的同时,只会增加14%的性能开销和19%的内存开销(对于占用空间较大的应用程序)。CachePerf检测到9个以前未知的错误。修复报告的bug可以实现从3%到3788%的性能加速。由于其有效性和低开销,CachePerf将成为现有分析器不可或缺的补充。
The cache plays a key role in determining the performance of applications, no matter for sequential or concurrent programs on homogeneous and heterogeneous architecture. Fixing cache misses requires to understand the origin and the type of cache misses. However, this remains to be an unresolved issue even after decades of research. This paper proposes a unified profiling tool--CachePerf--that could correctly identify different types of cache misses, differentiate allocator-induced issues from those of applications, and exclude minor issues without much performance impact. The core idea behind CachePerf is a hybrid sampling scheme: it employs the PMU-based coarse-grained sampling to select very few susceptible instructions (with frequent cache misses) and then employs the breakpoint-based fine-grained sampling to collect the memory access pattern of these instructions. Based on our evaluation, CachePerf only imposes 14% performance overhead and 19% memory overhead (for applications with large footprints), while identifying the types of cache misses correctly. CachePerf detected 9 previous-unknown bugs. Fixing the reported bugs achieves from 3% to 3788% performance speedup. CachePerf will be an indispensable complementary to existing profilers due to its effectiveness and low overhead.
Featherlight 即时错误共享检测
DOI: --
发表时间: 2018
期刊: ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子: --
作者:
Milind Chabbi;Shasha Wen;Xu Liu
通讯作者: Xu Liu
详细的缓存模拟,用于检测瓶颈、缺失原因和优化潜力
DOI: --
发表时间: 2006
期刊: ValueTools
影响因子: --
作者:
J. Tao;Wolfgang Karl
通讯作者: Wolfgang Karl
DMon:使用选择性分析有效检测和纠正数据局部性问题
DOI: --
发表时间: 2021
期刊: Proceedings of the Symposium on Operating Systems Principles
影响因子: --
作者:
Khan, Tanvir Ahmed;Neal, Ian;Pokam, Gilles;Mozafari, Barzan;Kasikci
通讯作者: Kasikci
DOI: 10.1109/cgo.2013.6495008
发表时间: 2013-02
期刊: Proceedings of the 2013 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子: --
作者:
Bin Bao;C. Ding
通讯作者: Bin Bao;C. Ding
用于数组收缩的集体循环融合
DOI: --
发表时间: 1992
期刊: International Workshop on Languages and Compilers for Parallel Computing
影响因子: --
作者:
G. Gao;R. Olsen;Vivek Sarkar;R. Thekkath
通讯作者: R. Thekkath