Design Space Exploration of Hybrid Ultra Low Power Branch Predictors

Design Space Exploration of Hybrid Ultra Low Power Branch Predictors
复制标题

DOI:
10.1007/978-3-642-28293-5_16
复制
发表时间:
2012-02
期刊:
--
影响因子:
--
通讯作者:
M. Bielby;Miles Gould;N. Topham
M. Bielby;Miles Gould;N. Topham
中科院分区:
其他
文献类型:
--
作者:
M. Bielby;Miles Gould;N. Topham

文献摘要

被引文献

相似文献

现代分支预测器通常太大且耗电,对于小型嵌入式处理器来说不是一个可行的选择,因为芯片空间、功耗和性能都非常宝贵。对于嵌入式处理器,高性能分支预测所需的大型高速缓存结构很容易占用比处理器其余部分组合更多的芯片空间。随着技术进步到 45nm 及以上,大泄漏能量将成为一个日益严重的问题,再加上大泄漏能量,根本不使用动态分支预测器通常看起来很有吸引力。本文试图找到一种在适合嵌入式处理器的混合预测器配置中使用超小型分支预测器的方法。我们引入了一种新颖的偏差参数来考虑何时静态或动态执行分支,进一步探索性能与能量的权衡。我们提出了一种解决方案,可以减少动态分支预测器混叠、提高性能并需要最少的额外芯片空间。给出的结果涉及芯片空间要求、能源使用和性能影响。我们研究如何以一种通常不被考虑的方式以及比之前提出的更低的比特预算来最好地优化这种平衡。 EEMBC 1.1 基准套件 [1] 用于探索能源与性能的权衡边界,取 31 个不同基准结果的平均值。我们通过使用 9 组新颖的偏差值,将 GShare 动态预测与分析的后向采取前向不采取 (BTFN)/后向不采取前向采取 (BNFT) 静态预测相结合,评估 5 种传统分支预测器配置和 36 个新型超小型混合分支预测器。结果表明,使用静态-动态混合不仅有益,而且对于非常小的预测器来说是必要的,可以对处理器的周期计数和总体能源使用产生积极影响。通过使用我们新颖的偏差参数,我们探索了性能与能量的权衡,并表明,通过对给定架构的峰值性能(500MHz 时 0.1 秒或 0.35%)的小幅降低(总运行时间在 28.35 秒范围内),我们可以通过减少动态预测器访问来获得显着的动态节能(最多消除传统混合预测器访问的 16.5%,即 5300 万次)。我们性能最佳的架构显示,与静态 BTFN 基线(总运行时间 30.46 秒)相比,运行时间平均缩短了 2 秒(6.7%),而成本仅增加了 0.01mm2(或 1%)的芯片空间。
Modern branch predictors are often too large and power hungry to be a viable option for small, embedded processors where die space, power consumption and performance are all at a premium. With embedded processors the large cache structures required for high performance branch prediction can easily take up more die space than the rest of the processor combined. When coupled with the large leakage energies, which are set to be an increasing issue as technologies advance to 45nm and beyond, it can often appear appealing to not use a dynamic branch predictor at all. This paper seeks to find a way of using an ultra small branch predictor in a hybrid predictor configuration suitable for an embedded processor. We introduce a novel bias parameter to the consideration of when to execute branches statically or dynamically, further exploring the performance vs energy trade-off. We present a solution that reduces dynamic branch predictor aliasing, improves performance and requires a minimum of extra die space. The results presented relate die space requirements, energy use and performance impacts. We look at how best to optimise this balance in a way that is usually not considered, and on a lower bits budget than has previously been presented. The EEMBC 1.1 benchmark suite [1] was used to explore the energy vs performance trade-off boundary, taking averages of the results across 31 different benchmarks. We evaluate 5 traditional branch predictor configurations and 36 novel ultra small hybrid branch predictors through the use of 9 sets of our novel bias values, combining GShare dynamic predictions with profiled backwards taken forwards not-taken (BTFN)/ backwards not-taken forwards taken (BNFT) static predictions. The results demonstrate that the use of a static-dynamic hybrid is not only beneficial but necessary for very small predictors to produce a positive effect on the cycle count and overall energy use of the processor. Through the use of our novel bias parameter we explore the performance vs energy trade-off and show that through a small (0.1 seconds at 500MHz or 0.35%) reduction in peak performance (total runtime in region of 28.35 seconds) for a given architecture we can gain substantial dynamic energy savings from reduced dynamic predictor accesses (removing up to an additional 16.5%, or 53 million, of the traditional hybrid predictor accesses). Our best performing architecture showed an average improvement in run time of 2 seconds (6.7%) over a static BTFN baseline (total runtime 30.46s), at the cost of only an additional 0.01mm2(or 1%) die space.