Beyond Static Parallel Loops: Supporting Dynamic Task Parallelism on Manycore Architectures with Software-Managed Scratchpad Memories

Beyond Static Parallel Loops: Supporting Dynamic Task Parallelism on Manycore Architectures with Software-Managed Scratchpad Memories
复制标题

DOI:
10.1145/3582016.3582020
复制
发表时间:
2023-03
期刊:
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3
影响因子:
--
通讯作者:
Lin Cheng;Max Ruttenberg;Dai Cheol Jung;D. Richmond;M. Taylor;M. Oskin;C. Batten
Lin Cheng;Max Ruttenberg;Dai Cheol Jung;D. Richmond;M. Taylor;M. Oskin;C. Batten
中科院分区:
其他
文献类型:
--
作者:
Lin Cheng;Max Ruttenberg;Dai Cheol Jung;D. Richmond;M. Taylor;M. Oskin;C. Batten

文献摘要

相似文献

许多核心体系结构通过使用简单的内核和简单的存储系统在单个芯片上集成了数百个核心数据移动和同步的正确性和性能。体系结构。同时,动态任务并行编程模型在解决多核处理器的可编程挑战方面,具有数十个复杂的核心和硬件cache连贯性。在大多数动态的任务并行编程模型中,在这项工作中不适合许多核心体系结构。使用SPM的体系结构,但是在执行这些架构时,我们还可以显着提高不规则工作负载的性能。 –28.5倍的工作负载加速,从我们的技术中受益,并且仅引起最小的开销,从而使没有的工作负载。
Manycore architectures integrate hundreds of cores on a single chip by using simple cores and simple memory systems usually based on software-managed scratchpad memories (SPMs). However, such architectures are notoriously challenging to program, since the programmers need to manually manage all aspects of data movement and synchronization for both correctness and performance. We argue that this manycore programmability challenge is one of the key barriers to achieving the promise of manycore architectures. At the same time, the dynamic task parallel programming model is enjoying considerable success in addressing the programmability challenge of multi-core processors with tens of complex cores and hardware cache coherence. Conventional wisdom suggests a work-stealing runtime, which forms the core of most dynamic task parallel programming models, is ill-suited for manycore architectures. In this work, we demonstrate that a work-stealing runtime is not just feasible on manycore architectures with SPMs, but such a runtime can also significantly improve the performance of irregular workloads when executing on these architectures. We also explore three optimizations that allow the runtime to leverage unused SPM space for further performance benefit. Our dynamic task parallel programming framework achieves 1.2–28.5× speedup on workloads that benefit from our techniques, and only induces minimal overhead for workloads that do not.