Control-Flow Decoupling

Control-Flow Decoupling
复制标题

控制流解耦

DOI:
10.1109/micro.2012.38
复制
发表时间:
2012
期刊:
2012 45th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
E. Rotenberg
E. Rotenberg
中科院分区:
--
文献类型:
--
作者:
Rami Sheikh;James Tuck;E. Rotenberg

文献摘要

被引文献

相似文献

移动和PC/服务器级处理器公司继续推出比其前身更快的旗舰核心微架构。同时,在芯片上放置更多的核,加上恒定的供电电压,使得每个核的能耗变得非常高。因此,我们面临的挑战是找到未来的微架构优化方案,既要提高性能,又要节约能源。消除分支的错误预测——既浪费时间又浪费精力——在这方面是有价值的。我们首先通过描述四个基准套件中的错误预测来探索控制流领域。我们发现三分之一的每k指令错误预测(MPKI)来自我们所谓的可分离分支:具有大控制依赖区域的分支(不适合if转换),其向后切片不依赖于它们的控制依赖指令或只有短依赖。我们提出了控制流解耦(CFD)来消除可分离分支的错误预测。其思想是将包含分支的循环分成两个循环:第一个循环仅包含分支的谓词计算,第二个循环包含分支及其控制相关指令。第一个循环通过体系结构队列将分支结果传递给第二个循环。从微观架构上讲,队列驻留在获取单元中,以驱动及时的、非推测性的获取或跳过控制依赖区域的连续动态实例。程序员或编译器都可以为CFD转换循环,我们对两者都求值。在类似英特尔Sandy Bridge核心的微架构上,CFD的性能提高了43%,能耗降低了41%。此外,对于某些应用程序,CFD是未来复杂有效的大窗口架构的必要催化剂,以容忍内存延迟。
Mobile and PC/server class processor companies continue to roll out flagship core micro architectures that are faster than their predecessors. Meanwhile placing more cores on a chip coupled with constant supply voltage puts per-core energy consumption at a premium. Hence, the challenge is to find future micro architecture optimizations that not only increase performance but also conserve energy. Eliminating branch mispredictions -- which waste both time and energy -- is valuable in this respect. We first explore the control-flow landscape by characterizing mispredictions in four benchmark suites. We find that a third of mispredictions-per-1K-instructions (MPKI) come from what we call separable branches: branches with large control-dependent regions (not suitable for if-conversion), whose backward slices do not depend on their control-dependent instructions or have only a short dependence. We propose control-flow decoupling (CFD) to eradicate mispredictions of separable branches. The idea is to separate the loop containing the branch into two loops: the first contains only the branch's predicate computation and the second contains the branch and its control-dependent instructions. The first loop communicates branch outcomes to the second loop through an architectural queue. Micro architecturally, the queue resides in the fetch unit to drive timely, non-speculative fetching or skipping of successive dynamic instances of the control-dependent region. Either the programmer or compiler can transform a loop for CFD, and we evaluate both. On a micro architecture configured similar to Intel's Sandy Bridge core, CFD increases performance by up to 43%, and reduces energy consumption by up to 41%. Moreover, for some applications, CFD is a necessary catalyst for future complexity-effective large-window architectures to tolerate memory latency.