Boomerang: A Metadata-Free Architecture for Control Flow Delivery

Boomerang: A Metadata-Free Architecture for Control Flow Delivery
复制标题

Boomerang:用于控制流交付的无元数据架构

DOI:
10.1109/hpca.2017.53
复制
发表时间:
2017
期刊:
2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
V. Nagarajan
V. Nagarajan
中科院分区:
--
文献类型:
--
作者:
Rakesh Kumar;Cheng;Boris Grot;V. Nagarajan

文献摘要

被引文献

相似文献

当代服务器工作负载的特点是来自深层分层软件堆栈的大量指令占用空间。整个堆栈的活动指令工作集可以轻松达到兆字节,从而导致由于指令缓存未命中而导致的频繁前端停顿和由于分支目标缓冲区(BTB)未命中而导致的流水线刷新。虽然已经提出了许多技术来解决这些问题,但它们中的每一个都需要专用的元数据结构,从而转化为显著的存储和复杂性成本。在本文中,我们问的问题是,是否有可能实现高性能的控制流交付,而无需元数据成本的现有技术。我们重新审视了以前提出的分支预测器定向预取方法,该方法仅利用分支预测器和BTB,通过在核心前端之前探索程序控制流来发现和预取丢失的指令缓存块。与传统的智慧相反,我们发现,这种方法可以有效地覆盖指令高速缓存未命中的现代CMP与长LLC访问延迟和多MB的服务器二进制文件。我们的第一个贡献在于解释分支预测器导向预取的功效的原因。我们的第二个贡献是Boomerang,一个用于控制流交付的无元数据架构。Boomerang利用分支预测器定向预取器来发现和预填充指令缓存块以及丢失的BTB条目。至关重要的是,我们证明,所需的额外硬件成本,以确定和填补BTB的失误是微不足道的。我们的实验评估表明,Boomerang匹配的性能的国家的最先进的控制流交付计划,而后者的高元数据和复杂性的开销。
Contemporary server workloads feature massive instruction footprints stemming from deep, layered software stacks. The active instruction working set of the entire stack can easily reach into megabytes, resulting in frequent front-end stalls due to instruction cache misses and pipeline flushes due to branch target buffer (BTB) misses. While a number of techniques have been proposed to address these problems, every one of them requires dedicated metadata structures, translating into significant storage and complexity costs. In this paper, we ask the question whether it is possible to achieve high-performance control flow delivery without the metadata costs of prior techniques. We revisit a previously proposed approach of branch-predictor-directed prefetching, which leverages just the branch predictor and BTB to discover and prefetch the missing instruction cache blocks by exploring the program control flow ahead of the core front-end. Contrary to conventional wisdom, we find that this approach can be effective in covering instruction cache misses in modern CMPs with long LLC access latencies and multi-MB server binaries. Our first contribution lies in explaining the reasons for the efficacy of branch-predictor-directed prefetching. Our second contribution is in Boomerang, a metadata-free architecture for control flow delivery. Boomerang leverages a branch-predictor-directed prefetcher to discover and prefill not only the instruction cache blocks, but also the missing BTB entries. Crucially, we demonstrate that the additional hardware cost required to identify and fill BTB misses is negligible. Our experimental evaluation shows that Boomerang matches the performance of the state-of-the-art control flow delivery scheme without the latter's high metadata and complexity overheads.