TSO_ATOMICITY: efficient hardware primitive for TSO-preserving region optimizations

TSO_ATOMICITY: efficient hardware primitive for TSO-preserving region optimizations
复制标题

TSO_ATOMICITY:用于 TSO 保留区域优化的高效硬件原语

DOI:
--
复制
发表时间:
2013
期刊:
International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Youfeng Wu
Youfeng Wu
中科院分区:
--
文献类型:
--
作者:
Cheng Wang;Youfeng Wu

文献摘要

被引文献

相似文献

基于数据依赖的程序优化可能无法保持程序中的内存一致性。以前的作品利用硬件原子原语来限制线程交织,以保持区域优化中的顺序一致性。然而,ATOMICITY原语对线程交错的限制过多,以优化使用流行的Total-Store-Ordering(TSO)内存一致性开发的实际应用程序,TSO内存一致性比顺序一致性弱。本文提出了一种新的硬件TSO_ATOMICITY原语,它对线程交错的限制比ATOMICITY原语少,程序执行效率比ATOMICITY原语高,但在所有区域优化中仍能保持TSO内存一致性。此外,TSO_ATOMICITY原语需要与ATOMICITY原语类似的体系结构支持,并且只需对现有ATOMICITY原语实现进行轻微更改即可实现。我们的实验结果表明,在一个最先进的动态二进制优化系统上的一个大的工作负载集,原子原语只能提高4%的平均性能。TSO_ATOMICITY原语可以减少与ATOMICITY原语相关的开销,并将性能平均提高12%。
Program optimizations based on data dependences may not preserve the memory consistency in the programs. Previous works leverage a hardware ATOMICITY primitive to restrict the thread interleaving for preserving sequential consistency in region optimizations. However, ATOMICITY primitive is over restrictive on the thread interleaving for optimizing real-world applications developed with the popular Total-Store-Ordering (TSO) memory consistency, which is weaker than sequential consistency. In this paper, we present a novel hardware TSO_ATOMICITY primitive, which has less restriction on the thread interleaving than ATOMICITY primitive to permit more efficient program execution than ATOMICITY primitive, but can still preserve TSO memory consistency in all region optimizations. Furthermore, TSO_ATOMICITY primitive requires similar architecture support as ATOMICITY primitive and can be implemented with only slight change to the existing ATOMICITY primitive implementation. Our experimental results show that in a start-of-art dynamic binary optimization system on a large set of workloads, ATOMICITY primitive can only improve the performance by 4% on average. TSO_ATOMICITY primitive can reduce the overhead associated with ATOMICITY primitive and improve the performance by 12% on average.