Configurable cache subsetting for fast cache tuning

Configurable cache subsetting for fast cache tuning
复制标题

可配置的缓存子集,用于快速缓存调整

DOI:
--
复制
发表时间:
2006
期刊:
2006 43rd ACM/IEEE Design Automation Conference
影响因子:
--
通讯作者:
F. Vahid
F. Vahid
中科院分区:
--
文献类型:
--
作者:
Pablo Viana;A. Gordon;Eamonn J. Keogh;E. Barros;F. Vahid

文献摘要

被引文献

相似文献

近年来,在商业微处理器中已经提出了具有诸如总大小、行大小和关联性的可变参数的多种可配置高速缓存的变体。研究表明,将可配置缓存调优到目标应用程序可以将内存访问功率降低50%以上。然而,在配置空间中搜索最佳配置可能需要大量时间或精力,即使在使用最新的高速缓存调优试探法时也是如此。我们试图为特定的应用程序域确定仍可实现有效调优的最小缓存配置子集。对于一套34个基准测试和具有18个可能配置的高速缓存,我们通过对所有可能子集的穷举搜索确定只需要3到4个候选配置来支持调优。我们引入了一种新的启发式算法,改编自为数据挖掘开发的高效启发式算法,以快速确定任意大小子集的最佳配置,并获得接近最优的结果。然后,我们考虑一个具有17,640种可能配置的可配置缓存,并改进我们的启发式方法,以包括预裁剪步骤,从而产生接近最佳的调优结果。我们的结论是,只需要3到4种可能的缓存配置,就可以为我们套件中的每个基准测试提供近乎最佳的配置-与最先进的缓存调整启发式方法相比,设计空间探索时间减少了91%
Numerous variations of configurable caches, having variable parameters like total size, line size, and associativity, have been proposed in commercial microprocessors in recent years. Tuning a configurable cache to a target application has been shown to reduce memory-access power by over 50%. However, searching the configuration space for the best configuration can require much time or power, even when using recent cache tuning heuristics. We sought to determine, for a particular domain of applications, the smallest subset of cache configurations that would still enable effective tuning. For a suite of 34 benchmarks and a cache with 18 possible configurations, we determine through an exhaustive search of all possible subsets, that only 3 or 4 candidate configurations are necessary to support tuning. We introduce a new heuristic, adapted from an efficient and effective heuristic developed for data mining, to quickly determine the best configurations for any sized subset, with near optimal results. We then consider a configurable cache with 17,640 possible configurations and improve our heuristic to include a pre-pruning step, yielding near optimal tuning results. We conclude that only 3 or 4 possible cache configurations are needed to offer a near optimal configuration for every benchmark in our suite - resulting in a 91% reduction in design space exploration time over a state-of-the-art cache tuning heuristic