Hardware Multithreaded Transactions

Hardware Multithreaded Transactions
复制标题

硬件多线程事务

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
David I. August
David I. August
中科院分区:
--
文献类型:
--
作者:
Jordan Fix;N. P. Nagendra;Sotiris Apostolakis;Hansen Zhang;Sophie Qiu;David I. August

文献摘要

被引文献

相似文献

事务内存系统中的推测有助于程序员和编译器生成高效的线程级并行程序。先前的研究表明,支持可跨多个线程的事务(而非要求事务必须在单个线程内)为程序员和并行化编译器提供了新型的推测并行化技术。不幸的是,对多线程事务(MTX)的软件支持会带来显著的额外线程间通信开销,用于推测验证。对于具有大量读写集合的程序,这种开销可能会使原本良好的并行化变得无利可图。一些使用先前软件MTX的程序通过专家程序员的大量努力克服了这一问题,他们尽量减小这些集合并优化通信,而编译器技术一直无法等效地实现这些能力。相反,本文通过低开销的推测验证使推测并行化更省力且更可行,提出了硬件MTX的首个完整设计、实现和评估。即使对包含数千万条指令的事务内的每一次加载和存储都进行最大限度的推测验证,也能实现复杂程序的高效并行化。在8个基准测试中,该系统在一个4核多核机器上相对于顺序执行实现了99%的几何平均加速比。
Speculation with transactional memory systems helps pro- grammers and compilers produce profitable thread-level parallel programs. Prior work shows that supporting transactions that can span multiple threads, rather than requiring transactions be contained within a single thread, enables new types of speculative parallelization techniques for both programmers and parallelizing compilers. Unfortunately, software support for multi-threaded transactions (MTXs) comes with significant additional inter-thread communication overhead for speculation validation. This overhead can make otherwise good parallelization unprofitable for programs with sizeable read and write sets. Some programs using these prior software MTXs overcame this problem through significant efforts by expert programmers to minimize these sets and optimize communication, capabilities which compiler technology has been unable to equivalently achieve. Instead, this paper makes speculative parallelization less laborious and more feasible through low-overhead speculation validation, presenting the first complete design, implementation, and evaluation of hardware MTXs. Even with maximal speculation validation of every load and store inside transactions of tens to hundreds of millions of instructions, profitable parallelization of complex programs can be achieved. Across 8 benchmarks, this system achieves a geomean speedup of 99% over sequential execution on a multicore machine with 4 cores.