Interprocedural strength reduction of critical sections in explicitly-parallel programs

Interprocedural strength reduction of critical sections in explicitly-parallel programs
复制标题

显式并行程序中关键部分的过程间强度降低

DOI:
--
复制
发表时间:
2013
期刊:
Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Vivek Sarkar
Vivek Sarkar
中科院分区:
--
文献类型:
--
作者:
R. Barik;Jisheng Zhao;Vivek Sarkar

文献摘要

被引文献

相似文献

在本文中,我们介绍了新颖的编译器优化技术,以减少在显式并行程序中发生的关键部分中执行的操作数量。具体而言,我们关注三个代码转换:1)关键部分的局部强度降低(PSR),以通过某些控制流动路径上的非临界部分替换临界部分; 2)临界负载消除(CLE)替换关键部分中的内存访问,访问临界部分之外加载的值的标量临时性访问; 3)非关键代码运动(NCM)从关键部分提起螺纹 - 本地计算。通过分解分析进一步提高了前两个转变的有效性。已经证明了来自三种不同显式并行编程模型的关键部分结构的技术的有效性 - Habanero Java(HJ)中的隔离结构,标准Java中的同步结构以及基于Java的Deuce软件交易软件系统的交易。我们使用了两个SMP平台(16核Intel Xeon SMP和一个32核IBM Power7 SMP)来评估我们对17个跨越所有三个模型的显式平行基准程序的优化。我们的结果表明,本文介绍的优化可以提供可衡量的性能改进,当该程序使用大量的处理器核心运行时,可以增加幅度。这些结果强调了优化关键部分的重要性,而这种优化的好处将随着未来多核处理器的核心数量的增加而继续增加。
In this paper, we introduce novel compiler optimization techniques to reduce the number of operations performed in critical sections that occur in explicitly-parallel programs. Specifically, we focus on three code transformations: 1) Partial Strength Reduction (PSR) of critical sections to replace critical sections by non-critical sections on certain control flow paths; 2) Critical Load Elimination (CLE) to replace memory accesses within a critical section by accesses to scalar temporaries that contain values loaded outside the critical section; and 3) Non-critical Code Motion (NCM) to hoist thread-local computations out of critical sections. The effectiveness of the first two transformations is further increased by interprocedural analysis. The effectiveness of our techniques has been demonstrated for critical section constructs from three different explicitly-parallel programming models - the isolated construct in Habanero Java (HJ), the synchronized construct in standard Java, and transactions in the Java-based Deuce software transactional memory system. We used two SMP platforms (a 16-core Intel Xeon SMP and a 32-Core IBM Power7 SMP) to evaluate our optimizations on 17 explicitly-parallel benchmark programs that span all three models. Our results show that the optimizations introduced in this paper can deliver measurable performance improvements that increase in magnitude when the program is run with a larger number of processor cores. These results underscore the importance of optimizing critical sections, and the fact that the benefits from such optimizations will continue to increase with increasing numbers of cores in future many-core processors.