Speculative Parallelization in Decoupled Look-ahead

Speculative Parallelization in Decoupled Look-ahead
复制标题

解耦前瞻中的推测并行化

DOI:
--
复制
发表时间:
2011
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Michael C. Huang
Michael C. Huang
中科院分区:
--
文献类型:
--
作者:
Alok Garg;Raj Parihar;Michael C. Huang

文献摘要

被引文献

相似文献

虽然规范的乱序引擎可以有效地利用顺序程序中的隐式并行性,但其有效性通常受到表现为分支误预测和高速缓存未命中的指令和数据供应缺陷的阻碍。由已执行程序的一部分引导的准确和深度的前瞻是一种简单而有效的方法,可以减轻分支预测错误和缓存未命中对性能的影响。不幸的是,程序片引导的前瞻往往受到前瞻代码片的速度的限制,特别是对于不规则的程序。在本文中,我们试图加快前瞻性代理使用推测并行化,这是特别适合的任务。首先,切片前瞻往往会减少重要的数据依赖,阻止成功的推测并行化。第二,前瞻任务不是正确性关键的,因此自然容忍依赖违反。这使得实现完全放弃了违规检测,极大地简化了架构支持。在一个简单的实现中,将推测并行化结合到前瞻代理中,进一步提高了系统性能,最高可达1.39倍,平均为1.13倍。
While a canonical out-of-order engine can effectively exploit implicit parallelism in sequential programs, its effectiveness is often hindered by instruction and data supply imperfections manifested as branch mispredictions and cache misses. Accurate and deep look-ahead guided by a slice of the executed program is a simple yet effective approach to mitigate the performance impact of branch mispredictions and cache misses. Unfortunately, program slice-guided look ahead is often limited by the speed of the look-ahead code slice, especially for irregular programs. In this paper, we attempt to speed up the look-ahead agent using speculative parallelization, which is especially suited for the task. First, slicing for look-ahead tends to reduce important data dependences that prohibit successful speculative parallelization. Second, the task for look-ahead is not correctness critical and thus naturally tolerates dependence violations. This enables an implementation to forgo violation detection altogether, simplifying architectural support tremendously. In a straightforward implementation, incorporating speculative parallelization to the look-ahead agent further improves system performance by up to 1.39x with an average of 1.13x.