FastLane: improving performance of software transactional memory for low thread counts

FastLane: improving performance of software transactional memory for low thread counts
复制标题

FastLane:提高软件事务内存的性能以实现低线程数

DOI:
10.1145/2442516.2442528
复制
发表时间:
2013
期刊:
影响因子:
0.9
通讯作者:
Gilles Muller
Gilles Muller
中科院分区:
数学3区
文献类型:
--
作者:
Jons;C. Fetzer;P. Felber;E. Rivière;Gilles Muller

文献摘要

被引文献

相似文献

软件事务存储器(STM)可以实现可伸缩的并发程序实现,因为应用程序的相对性能随着支持它的线程数的增加而增加,但是,绝对性能通常会受到事务管理和对共享存储器的插装访问的开销的影响。这通常会导致线程数较低的基于STM的程序的性能比同一应用程序的顺序、非检测版本差。 在本文中,我们提出了FastLane,一个新的STM算法,桥梁之间的性能差距顺序执行和经典的STM算法时,运行在几个核心。FastLane寻求降低仪器成本,从而降低其目标操作范围内的性能下降。我们引入了一种新的算法,区分两种类型的线程:一个线程(主)悲观地执行事务,从来没有中止,从而以最小的仪器和管理成本,而其他线程(助手)可以提交投机交易时,他们不与主冲突。因此,帮助程序有助于应用程序的进展,而不会损害主程序的性能。 我们实现FastLane作为一个国家的最先进的STM运行时系统和编译器的扩展。产生多个代码路径以在单个、几个和多个核心上执行。运行时系统根据目标机器上可用的内核数量选择提供最佳吞吐量的代码路径。评估结果表明,我们的方法在低线程数下提供了有前途的性能:FastLane几乎系统地在1-6个线程范围内战胜了经典STM,并且通常比从2个线程开始的相同应用程序的非仪表化版本的顺序执行更好。
Software transactional memory (STM) can lead to scalable implementations of concurrent programs, as the relative performance of an application increases with the number of threads that support it. However, the absolute performance is typically impaired by the overheads of transaction management and instrumented accesses to shared memory. This often leads STM-based programs with low thread counts to perform worse than a sequential, non-instrumented version of the same application. In this paper, we propose FastLane, a new STM algorithm that bridges the performance gap between sequential execution and classical STM algorithms when running on few cores. FastLane seeks to reduce instrumentation costs and thus performance degradation in its target operation range. We introduce a novel algorithm that differentiates between two types of threads: One thread (the master) executes transactions pessimistically without ever aborting, thus with minimal instrumentation and management costs, while other threads (the helpers) can commit speculative transactions only when they do not conflict with the master. Helpers thus contribute to the application progress without impairing on the performance of the master. We implement FastLane as an extension of a state-of-the-art STM runtime system and compiler. Multiple code paths are produced for execution on a single, few, and many cores. The runtime system selects the code path providing the best throughput, depending on the number of cores available on the target machine. Evaluation results indicate that our approach provides promising performance at low thread counts: FastLane almost systematically wins over a classical STM in the 1-6 threads range, and often performs better than sequential execution of the non-instrumented version of the same application starting with 2 threads.