Compiler estimation of load imbalance overhead in speculative parallelization

Compiler estimation of load imbalance overhead in speculative parallelization
复制标题

推测并行化中负载不平衡开销的编译器估计

DOI:
--
复制
发表时间:
2004
期刊:
Proceedings. 13th International Conference on Parallel Architecture and Compilation Techniques, 2004. PACT 2004.
影响因子:
--
通讯作者:
Marcelo H. Cintra
Marcelo H. Cintra
中科院分区:
--
文献类型:
--
作者:
J. Dou;Marcelo H. Cintra

文献摘要

被引文献

相似文献

推测并行化是一种补充自动编译器并行化的技术,它允许编译器无法完全分析的代码段主动并行执行。然而,虽然推测并行化可能会带来显着的加速,但与该技术相关的一些开销在实践中限制了这些加速。我们提出了一种新颖的推测性多线程执行编译器模型,可用于推断推测性并行化的开销和预期的性能增益或损失。该模型基于估计线程的可能执行持续时间,适当地考虑了大多数推测执行环境的调度限制,并且可以包括所有推测并行化开销。此外,与尝试定性估计推测多线程执行的潜在“好”或“坏”部分的启发法不同,该模型允许编译器定量估计加速或减速。然后,编译器或运行时系统可以使用这种定量估计来做出更复杂和更有根据的权衡决策。我们在一系列 SPEC 基准测试中的多个循环上使用了所提出的框架,这些循环主要受到负载不平衡以及线程分派和提交开销的影响。实验结果表明,我们的框架平均可以识别 68% 的导致速度减慢的循环,平均 97% 的导致加速的循环。事实上,我们的框架对基准测试中平均 44% 的循环的加速或减速预测误差小于 20%,对于 84% 的循环平均误差小于 50%。总体而言,与尝试推测性并行化所有考虑的循环的简单方法相比,我们的框架的性能平均提高了 5%,最高可达 38%。
Speculative parallelization is a technique that complements automatic compiler parallelization by allowing code sections that cannot be fully analyzed by the compiler to be aggressively executed in parallel. However, while speculative parallelization can potentially deliver significant speedups, several overheads associated with the technique limit these speedups in practice. We propose a novel compiler model of speculative multithreaded execution that can be used to reason about the overheads and expected resulting performance gains, or losses, from speculative parallelization. This model is based on estimating the likely execution duration of threads, properly takes into account the scheduling restrictions of most speculative execution environments, and can include all speculative parallelization overheads. Also, different from heuristics that attempt to qualitatively estimate potentially "good" or "bad" sections for speculative multithreaded execution, this model allows the compiler to estimate the speedup or slowdown quantitatively. Such quantitative estimate can then be used by the compiler or run-time system to make more complex and educated tradeoff decisions. We use the proposed framework on a number of loops from a collection of SPEC benchmarks that suffer mainly from load imbalance and thread dispatch and commit overheads. Experimental results show that our framework can identify on average 68% of the loops that cause slowdowns and on average 97% of the loops that lead to speedups. In fact, our framework predicts the speedups or slowdowns with an error of less than 20% for an average of 44% of the loops across the benchmarks, and with an error of less than 50% for an average of 84% of the loops. Overall, our framework leads to a performance improvement of 5% on average, and as high as 38%, over a naive approach that attempts to speculatively parallelize all the loops considered.