PDede: Partitioned, Deduplicated, Delta Branch Target Buffer

PDede: Partitioned, Deduplicated, Delta Branch Target Buffer
复制标题

DOI:
10.1145/3466752.3480046
复制
发表时间:
2021-10
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
N. Soundararajan;Peter Braun;Tanvir Ahmed Khan;Baris Kasikci;Heiner Litz;S. Subramoney
N. Soundararajan;Peter Braun;Tanvir Ahmed Khan;Baris Kasikci;Heiner Litz;S. Subramoney
中科院分区:
其他
文献类型:
--
作者:
N. Soundararajan;Peter Braun;Tanvir Ahmed Khan;Baris Kasikci;Heiner Litz;S. Subramoney

文献摘要

被引文献

相似文献

由于庞大的脚印,尽管对这些摊位有重要的贡献,但现代数据中心的损失经常遭受前端摊位通过更有效的替换政策和预摘要政策来提高BTB,以优化BTB的存储在这项工作中,效率不足。分支指令的数量具有相同的分支目标,(2)大量分支目标共享相同的页面地址,(3)很大一部分分支指令及其目标位于同一页面上,我们观察到,当应用程序的地址空间稀少,它们在页面内和跨页面中揭示了空间位置。基于这些见解的区域数量。分支机构之间的冗余及其目标。(a)BTB分区,(b)分支目标重复数据删除,以及(c)Delta分支目标编码,以减少BTB失误的前端档位。几种用法方案,并表明它通过将BTB的遗体平均减少54.7%(最高到最高,最高可提供14.4%(最高76%)IPC加速99.8%)。
Due to large instruction footprints, contemporary data center applications suffer from frequent frontend stalls. Despite being a significant contributor to these stalls, the Branch Target Buffer (BTB) has received less attention compared to other frontend structures such as the instruction cache. While prior works have looked at enhancing the BTB through more efficient replacement policies and prefetching policies, a thorough analysis into optimizing the BTB’s storage efficiency is missing. In this work, we analyze BTB accesses for a large number (100+) of frontend bound applications to understand their branch target characteristics. This analysis, provides three significant observations about the nature of branch targets: (1) a significant number of branch instructions have the same branch target, (2) a significant number of branch targets share the same page address, and (3) a significant percentage of branch instructions and their targets are located on the same page. Furthermore, we observe that while applications’ address spaces are sparsely populated, they exhibit spatial locality within and across pages. We refer to these multi-page addresses as regions and we show that applications traverse a significantly smaller number of regions than pages. Based on these insights, we propose PDede, an efficient re-design of the BTB micro-architecture that improves storage efficiency by removing redundancy among branches and their targets. PDede introduces three techniques, (a) BTB Partitioning, (b) Branch Target Deduplication, and (c) Delta Branch Target Encoding to reduce BTB miss induced frontend stalls. We evaluate PDede across 100+ applications, spanning several usage scenarios, and show that it provides an average 14.4% (up to 76%) IPC speedup by reducing BTB misses by 54.7% on average (and up to 99.8%).