Optimizing Work Stealing Communication with Structured Atomic Operations

Optimizing Work Stealing Communication with Structured Atomic Operations
复制标题

通过结构化原子操作优化工作窃取通信

DOI:
10.1145/3472456.3472522
复制
发表时间:
2021
期刊:
50th International Conference on Parallel Processing (ICPP '21
影响因子:
--
通讯作者:
Larkins, D. B.
Larkins, D. B.
中科院分区:
--
文献类型:
--
作者:
Cartier, H.;Dinan, J.;Larkins, D. B.

文献摘要

参考文献

被引文献

相似文献

依赖于稀疏或不规则数据的应用程序在现代分布式内存系统上进行扩展通常具有挑战性。因此,为了保持效率,这些系统通常需要持续的负载平衡。偷工作是一种补救失衡的常用方法。在这项工作中,我们提出了一种工作窃取策略,可以将窃取操作所需的通信量减少一半。我们展示了,作为管理本地队列状态的少量额外复杂性的交换,我们可以将发现和声明工作合并到单个步骤中。一般来说,窃取作品需要两个步骤:先发现作品,然后认领。我们的系统SWS提供了一种机制,其中两个进程在单一通信中执行,而不需要多个同步消息。通过原子操作的新应用程序可以减少通信,原子操作可以操作任务队列元数据的紧凑表示。我们使用测试动态负载平衡系统和执行不平衡树搜索的已知基准来证明该策略的有效性。我们的研究结果表明,通信的减少减少了任务获取时间和窃取时间,从而提高了稀疏计算的整体性能。
Applications that rely on sparse or irregular data are often challenging to scale on modern distributed-memory systems. As a result, these systems typically require continuous load balancing in order to maintain efficiency. Work stealing is a common technique to remedy imbalance. In this work we present a strategy for work stealing that reduces the amount of communication required for a steal operation by half. We show that in exchange for a small amount of additional complexity to manage the local queue state we can combine both discovering and claiming work into a single step. Conventionally, work stealing uses a two step process of discovering work and then claiming it. Our system, SWS, provides a mechanism where both processes are performed in a singular communication without the need for multiple synchronization messages. This reduction in communication is possible with the novel application of atomic operations that manipulate a compact representation of task queue metadata. We demonstrate the effectiveness of this strategy using known benchmarks for testing dynamic load balancing systems and for performing unbalanced tree searches. Our results show the reduction in communication reduces task acquisition time and steal time, which in turn improves overall performance on sparse computations.
自适应且可靠的并行计算9 工作站网络
DOI: --
发表时间: 1997
期刊:
影响因子: --
作者:
P. Lisiecki;R. Blumofe
通讯作者: R. Blumofe
优化的分布式工作窃取
DOI: --
发表时间: 2016
期刊: Workshop on Irregular Applications: Architectures and Algorithms
影响因子: --
作者:
Vivek Kumar;K. Murthy;Vivek Sarkar;Yili Zheng
通讯作者: Yili Zheng
开放社区运行时:用于超大规模计算的运行时系统
DOI: --
发表时间: 2016
期刊: IEEE Conference on High Performance Extreme Computing
影响因子: --
作者:
T. Mattson;R. Cledat;Vincent Cavé;Vivek Sarkar;Zoran Budimlic;S. Chatterjee;J. Fryman;Ivan B. Ganev;Rob C. Knauerhase;Min Lee;Benoît Meister;Brian R. Nickerson;Nick Pepperling;B. Seshasayee;Sagnak Tasirlar;J. Teller;Nick Vrvilo
通讯作者: Nick Vrvilo
NESL 的可证明时间和空间高效的实现
DOI: 10.1145/232627.232650
发表时间: 1996
期刊: Proceedings of the 19th International Symposium on Principles and Practice of Declarative Programming
影响因子: --
作者:
G. Blelloch;John Greiner
通讯作者: John Greiner
在众核上使用 Qthread 调度 Chapel 任务:两个调度程序的故事
DOI: --
发表时间: 2017
期刊: ROSS@HPDC
影响因子: --
作者:
N. Evans;Stephen L. Olivier;R. Barrett;George Stelle
通讯作者: George Stelle