Optimising power efficiency in trace cache fetch unit

Optimising power efficiency in trace cache fetch unit
复制标题

优化跟踪缓存获取单元的电源效率

DOI:
10.1049/iet-cdt:20060170
复制
发表时间:
2007
期刊:
IET Comput. Digit. Tech.
影响因子:
--
通讯作者:
M. Kandemir
M. Kandemir
中科院分区:
--
文献类型:
--
作者:
Jie S. Hu;N. Vijaykrishnan;M. J. Irwin;M. Kandemir

文献摘要

参考文献

被引文献

相似文献

随着超标量处理器的发布宽度和功能单元数量不断增加,取指单元必须支持大的取指带宽,以充分利用数据路径资源。由于传统的指令提取机制并未针对功耗进行优化,因此这种趋势使得提取单元中的功耗问题变得更加严重。本文探讨了传统指令缓存中由于动态控制流而导致的额外功耗问题。跟踪缓存在执行过程中捕获代码的动态路径/特征,为获取单元中的功耗优化提供了潜在的框架。我们的研究表明,传统跟踪缓存(CTC)由于同时访问跟踪缓存和指令缓存,可能会增加取指单元的功耗,而顺序跟踪缓存(STC)具有较低功耗的优势,但代价是显着的性能损失。为了解决这个问题,我们对跟踪分布和访问局部性进行了详细的研究。基于这项研究,我们首先提出了一种新模型:选择性跟踪缓存(SLTC)。 SLTC 使用编译器和硬件支持来有选择地控制跟踪缓存查找和更新。实验评估表明,我们的选择性跟踪缓存平均比 CTC 降低了 42.2% 的功耗,比 STC 额外降低了 21.8%,而与 CTC 相比,性能损失不超过 1.8%。此外,我们提出了一种基于动态方向预测的跟踪缓存(DPTC),它消除了 SLTC 中涉及的编译和指令集架构(ISA)修改的需要。在提取方向预测器的支持下,DPTC 实现了具有竞争力的电源效率。平均而言,与 CTC 和 STC 相比,DPTC 的取指单元功耗分别降低了 40.5% 和 17.6%,而 CTC 的性能损失不到 2.4%。
As the issue width and the number of function units of superscalar processors continue to increase, the fetch unit must support a large fetch bandwidth in order to fully utilise the datapath resources. This trend makes power issue worse in the fetch unit since the traditional instruction fetch mechanism is not optimised for power consumption. This paper explores the problem of extra power consumption in traditional instruction caches because of dynamic control flows. Capturing the dynamic paths/characteristics of code during the course of execution, trace caches provide a potential framework for power optimisation in the fetch unit. Our study shows that conventional trace caches (CTC) may increase power consumption in the fetch unit because of the simultaneous access to both the trace cache and the instruction cache, and sequential trace caches (STC) have the advantage of lower power consumption at the cost of a significant performance loss. In order to address this problem, we perform a detailed study of trace distribution and access locality. Based on this study, we first propose a new model, the selective trace cache (SLTC). SLTC uses both compiler and hardware support to selectively control trace cache lookup and update. Experimental evaluation shows that our selective trace cache achieves up to 42.2% power reduction over CTC and an additional reduction of up to 21.8% over STC, on the average, while only trading a performance loss of no more than 1.8% compared to CTC. Further, we propose a dynamic direction prediction based trace cache (DPTC), which eliminates the need for compilation and instruction set architecture (ISA) modification involved in SLTC. Powered by a fetch direction predictor, DPTC achieves competitive power efficiency. On the average, DPTC reduces the power consumption by up to 40.5% and 17.6% in the fetch unit compared to CTC and STC, respectively, by trading a performance loss of less than 2.4% to CTC.
DOI: --
发表时间: 2023
期刊: --
影响因子: --
作者:
D. Brooks;V. Tiwari;Intel Margaret Martonosi
通讯作者: D. Brooks;V. Tiwari;Intel Margaret Martonosi