Defensive loop tiling for shared cache

Defensive loop tiling for shared cache
复制标题

DOI:
10.1109/cgo.2013.6495008
复制
发表时间:
2013-02
期刊:
Proceedings of the 2013 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Bin Bao;C. Ding
Bin Bao;C. Ding
中科院分区:
其他
文献类型:
--
作者:
Bin Bao;C. Ding

文献摘要

被引文献

相似文献

循环平铺是一种编译器转换,它调整应用程序的工作集以适应缓存层次结构。在今天的多核处理器上,层次结构的一部分,特别是最后一级缓存(LLC)是共享的。共享缓存中的可用缓存空间根据共同运行的应用程序而变化。此外,在包含缓存层次结构的机器上,共享缓存中的干扰可能导致私有缓存中的驱逐,这个问题被称为包含受害者。本文介绍了防御性平铺,这是一套编译器技术,用于估计缓存共享的效果,然后选择可以在协同运行环境中提供健壮性能的平铺大小。转换的目标是优化缓存的使用,同时防止干扰。它完全是一种静态技术,不需要程序分析。本文展示了如何将其集成到生产质量的编译器中,并通过在实际系统上进行模拟和测试,评估了它对程序协同运行和单独运行性能的一组tililing基准的影响。
Loop tiling is a compiler transformation that tailors an application's working set to fit in a cache hierarchy. On today's multicore processors, part of the hierarchy especially the last level cache (LLC) is shared. The available cache space in shared cache changes depending on co-run applications. Furthermore on machines with an inclusive cache hierarchy, the interference in the shared cache can cause evictions in the private cache, a problem known as the inclusion victims. This paper presents defensive tiling, a set of compiler techniques to estimate the effect of cache sharing and then choose the tile sizes that can provide robust performance in co-run environments. The goal of the transformation is to optimize the use of the cache while at the same time guarding against interference. It is entirely a static technique and does not require program profiling. The paper shows how it can be integrated into a production-quality compiler and evalutes its effect on a set of tililing benchmarks for both program co-run and solo-run performance, using both simulation and testing on real systems.