Exploiting Concurrent GPU Operations for Efficient Work Stealing on Multi-GPUs

Exploiting Concurrent GPU Operations for Efficient Work Stealing on Multi-GPUs
复制标题

利用并发 GPU 操作在多 GPU 上高效窃取工作

DOI:
10.1109/sbac-pad.2012.28
复制
发表时间:
2012
期刊:
2012 IEEE 24th International Symposium on Computer Architecture and High Performance Computing
影响因子:
--
通讯作者:
Vincent Danjean
Vincent Danjean
中科院分区:
--
文献类型:
--
作者:
J. F. Lima;T. Gautier;N. Maillard;Vincent Danjean

文献摘要

被引文献

相似文献

Exascale计算的竞争自然导致当前技术融合到多CPU/多GPU计算机,基于通过PCI-Express总线或互连网络互连的数千个CPU和GPU。为了利用这种高计算能力,程序员必须解决在混合体系结构上调度并行程序的问题。而且,由于GPU的性能以比PCI总线的吞吐量快得多的速率增加,因此调度器必须有效地管理数据传输。本文针对多GPU计算节点,其中多个GPU连接到同一台机器。为了克服这种平台上的数据传输限制,可用的软件通常在执行之前计算任务的映射,该映射尊重它们的依赖性并最小化全局数据传输。这样的方法过于死板,不能使执行适应系统的可能变化或应用程序的负载。我们提出了一个解决方案,这是正交的上述:Xkaapi软件栈的扩展,使充分利用多GPU系统的性能,通过异步GPU任务。Xkaapi通过使用标准的工作窃取算法来调度任务,运行时可以有效地利用并发的GPU操作。运行时扩展使得在当前一代GPU上重叠数据传输和任务执行成为可能。我们证明,重叠的能力是至少一样重要的计算调度决策,以减少完成时间的并行程序。通过对两个稠密线性代数问题(矩阵积和Cholesky分解)的实验表明,该算法与其他基于静态调度的软件相比具有很强的竞争力。此外,我们能够保持最佳性能(约。310 GFlop/s),即使是不能完全存储在一个GPU内存中的矩阵。使用八个GPU,我们相对于单GPU实现了6.74的速度提升。我们的Cholesky分解的性能,任务之间的依赖关系更复杂,优于最先进的单GPU MAGMA代码。
The race for Exascale computing has naturally led the current technologies to converge to multi-CPU/multi-GPU computers, based on thousands of CPUs and GPUs interconnected by PCI-Express buses or interconnection networks. To exploit this high computing power, programmers have to solve the issue of scheduling parallel programs on hybrid architectures. And, since the performance of a GPU increases at a much faster rate than the throughput of a PCI bus, data transfers must be managed efficiently by the scheduler. This paper targets multi-GPU compute nodes, where several GPUs are connected to the same machine. To overcome the data transfer limitations on such platforms, the available soft wares compute, usually before the execution, a mapping of the tasks that respects their dependencies and minimizes the global data transfers. Such an approach is too rigid and it cannot adapt the execution to possible variations of the system or to the application's load. We propose a solution that is orthogonal to the above mentioned: extensions of the Xkaapi software stack that enable to exploit full performance of a multi-GPUs system through asynchronous GPU tasks. Xkaapi schedules tasks by using a standard Work Stealing algorithm and the runtime efficiently exploits concurrent GPU operations. The runtime extensions make it possible to overlap the data transfers and the task executions on current generation of GPUs. We demonstrate that the overlapping capability is at least as important as computing a scheduling decision to reduce completion time of a parallel program. Our experiments on two dense linear algebra problems (Matrix Product and Cholesky factorization) show that our solution is highly competitive with other soft wares based on static scheduling. Moreover, we are able to sustain the peak performance (approx. 310 GFlop/s) on DGEMM, even for matrices that cannot be stored entirely in one GPU memory. With eight GPUs, we archive a speed-up of 6.74 with respect to single-GPU. The performance of our Cholesky factorization, with more complex dependencies between tasks, outperforms the state of the art single-GPU MAGMA code.