TurboBŁYSK: Scheduling for Improved Data-Driven Task Performance with Fast Dependency Resolution
TurboBŁYSK: Scheduling for Improved Data-Driven Task Performance with Fast Dependency Resolution
复制标题
TurboBŁYSK:通过快速依赖性解析来提高数据驱动任务性能的调度
DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Vladimir Vlassov
中科院分区:
文献类型:
--
作者:
Artur Podobas;M. Brorsson;Vladimir Vlassov
Data-driven task-parallelism is attracting growing interest and has now been added to OpenMP (4.0). This paradigm simplifies the writing of parallel applications, extracting parallelism, and facilitates the use of distributed memory architectures. While the programming model itself is becoming mature, a problem with current run-time scheduler implementations is that they require a very large task granularity in order to scale. This limitation goes at odds with the idea of task-parallel programing where programmers should be able to concentrate on exposing parallelism with little regard to the task granularity. To mitigate this limitation, we have designed and implemented TurboBŁYSK, a highly efficient run-time scheduler of tasks with explicit data-dependence annotations. We propose a novel mechanism based on pattern-saving that allows the scheduler to re-use previously resolved dependency patterns, based on programmer annotations, enabling programs to use even the smallest of tasks and scale well. We experimentally show that our techniques in TurboBŁYSK enable achieving nearly twice the peak performance compared with other run-time schedulers. Our techniques are not OpenMP specific and can be implemented in other task-parallel frameworks.