Reducing the Power and Complexity of Path-Based Neural Branch Prediction

Reducing the Power and Complexity of Path-Based Neural Branch Prediction
复制标题

降低基于路径的神经分支预测的能力和复杂性

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Daniel A. Jiménez
Daniel A. Jiménez
中科院分区:
--
文献类型:
--
作者:
G. Loh;Daniel A. Jiménez

文献摘要

被引文献

相似文献

传统的基于路径的神经预测器(PBNP)实现了非常高的预测精度,但其非常深的流水线实现使其成为复杂且功耗密集的组件。大的复杂度和功率的主要原因之一是,对于h的历史长度,PBNP必须使用在anh级预测器流水线中组织的单独索引的SRAM阵列(或遭受非常长的更新延迟)。每个流水线级需要一个单独的行解码器,用于对应的SRAM阵列、级间锁存器、控制逻辑和检查点支持。所有这些都增加了预测器的功能和复杂性。我们提出了两种技术来解决这个问题。第一种是模路径历史,其将分支结果历史长度与路径历史长度相乘,从而允许更短的路径历史(并且因此更少的预测器流水线阶段),同时利用传统的长分支结果历史。流水线长度的减少导致降低的功率和实现复杂度。第二种技术是基于偏置的过滤(BBF),它利用了神经预测器已经有办法跟踪强偏置分支的事实。BBF使用偏置权重来过滤掉大部分总是被采用或大部分总是不被采用的分支,并避免消耗这些分支的更新功率。我们的建议是复杂有效的,因为它降低了功率和复杂性的PBNP没有负面影响的性能。模路径历史和BBF的组合导致32 KB和64 KB预测器的预测器准确度略微提高1%,但更重要的是,该技术通过将SRAM阵列的数量从30+减少到仅4-6个表并将预测器更新活动减少4- 5%来降低功率和复杂度。
A conventional path-based neural predictor (PBNP) achieves very high prediction accuracy, but its very deeply pipelined implementation makes it both a complex and power-intensive component. One of the major reasons for the large complexity and power is that for a history length of h, the PBNP must useh separately indexed SRAM arrays (or suffer from a very long update latency) organized in anh-stage predictor pipeline. Each pipeline stage requires a separate row-decoder for the corresponding SRAM array, inter-stage latches, control logic, and checkpointing support. All of these add power and complexity to the predictor. We propose two techniques to address this problem. The first is modulo path-historywhich decouples the branch outcome history length from the path history length allowing for a shorter path history (and therefore fewer predictor pipeline stages) while simultaneously making use of a traditional long branch outcome history. The pipeline length reduction results in decreased power and implementation complexity. The second technique is bias-based filtering(BBF) which takes advantage of the fact that neural predictors already have a way to track strongly biased branches. BBF uses the bias weights to filter out mostly always taken or mostly always not-taken branches and avoids consuming update power for such branches. Our proposal is complexity effective because it decreases the power and complexity of the PBNP without negatively impacting performance. The combination of modulo path-history and BBF results in a slight improvement in predictor accuracy of 1% for 32KB and 64KB predictors, but more importantly the techniques reduce power and complexity by reducing the number of SRAM arrays from 30+ down to only 4-6 tables, and reducing predictor update activity by 4-5%.