Exploiting dynamic timing margins in microprocessors for frequency-over-scaling with instruction-based clock adjustment

Exploiting dynamic timing margins in microprocessors for frequency-over-scaling with instruction-based clock adjustment
复制标题

利用微处理器中的动态时序裕度通过基于指令的时钟调整来实现超频

DOI:
--
复制
发表时间:
2015
期刊:
Design, Automation and Test in Europe
影响因子:
--
通讯作者:
A. Burg
A. Burg
中科院分区:
--
文献类型:
--
作者:
J. Constantin;Lai Wang;G. Karakonstantis;A. Chattopadhyay;A. Burg

文献摘要

被引文献

相似文献

静态时序分析为基于最坏情况关键路径设置微处理器内核的时钟周期提供了依据。然而,根据设计的不同,这个关键路径并不总是被激发的,因此理论上可以利用动态时间裕度来获得更好的速度或更低的功耗(通过电压缩放)。本文介绍了一种基于预测指令的动态时钟调整技术,以减少流水线微处理器的动态时间余量。为此,我们在动态变化的程序执行流程中利用单个指令的不同时序要求,而不需要复杂的电路级措施来检测和纠正时序违规。我们提供了一个设计流程,利用布局后动态时序分析来提取设计的动态时序信息,并将结果集成到定制的周期精确模拟器中。该模拟器允许注释单个指令及其对计时的影响(在每个管道阶段),并快速导出复杂基准测试的总体代码执行时间。设计方法在微架构层面进行了说明,展示了在28纳米CMOS技术的6级OpenRISC顺序通用处理器核心上可能获得的性能和功率增益。我们表明,与传统的同步时钟相比,采用指令依赖的动态时钟调整平均可以提高38%的运行速度或降低24%的功耗,而传统的同步时钟在任何时候都必须尊重通过静态定时分析确定的最坏时间。
Static timing analysis provides the basis for setting the clock period of a microprocessor core, based on its worst-case critical path. However, depending on the design, this critical path is not always excited and therefore dynamic timing margins exist that can theoretically be exploited for the benefit of better speed or lower power consumption (through voltage scaling). This paper introduces predictive instruction-based dynamic clock adjustment as a technique to trim dynamic timing margins in pipelined microprocessors. To this end, we exploit the different timing requirements for individual instructions during the dynamically varying program execution flow without the need for complex circuit-level measures to detect and correct timing violations. We provide a design flow to extract the dynamic timing information for the design using post-layout dynamic timing analysis and we integrate the results into a custom cycle-accurate simulator. This simulator allows annotation of individual instructions with their impact on timing (in each pipeline stage) and rapidly derives the overall code execution time for complex benchmarks. The design methodology is illustrated at the microarchitecture level, demonstrating the performance and power gains possible on a 6-stage OpenRISC in-order general purpose processor core in a 28nm CMOS technology. We show that employing instruction-dependent dynamic clock adjustment leads on average to an increase in operating speed by 38% or to a reduction in power consumption by 24%, compared to traditional synchronous clocking, which at all times has to respect the worst-case timing identified through static timing analysis.