Performance Portability of a GPU Enabled Factorization with the DAGuE Framework

Performance Portability of a GPU Enabled Factorization with the DAGuE Framework
复制标题

使用 DAGuE 框架支持 GPU 分解的性能可移植性

DOI:
10.1109/cluster.2011.51
复制
发表时间:
2011
期刊:
2011 IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
Jack J. Dongarra
Jack J. Dongarra
中科院分区:
--
文献类型:
--
作者:
G. Bosilca;A. Bouteiller;T. Hérault;P. Lemarinier;Narapat Ohm Saengpatsa;S. Tomov;Jack J. Dongarra

文献摘要

被引文献

相似文献

性能可移植性是当今开发人员在异类高性能计算机上面临的主要挑战,这些计算机包括互连、非统一访问的内存、多核和像GPU这样的加速器。最近的研究已经成功地证明,密集线性代数运算可以由运行时系统使用DAG表示来有效地处理。在这项工作中,我们介绍了DAGY运行时的GPU子系统,并在Cholesky因式分解测试用例上评估了程序员在DAGY框架中启用GPU加速所需的最小努力。这种不变的代码在各种不同的和分布式的许多内核和GPU资源上实现的性能证明了所需的性能可移植性。
Performance portability is a major challenge faced today by developers on heterogeneous high performance computers, consisting of an interconnect, memory with non-uniform access, many-cores and accelerators like GPUs. Recent studies have successfully demonstrated that dense linear algebra operations can be efficiently handled by runtime systems using a DAG representation. In this work, we present the GPU subsystem of the DAGuE runtime, and assess, on the Cholesky factorization test case, the minimal efforts required by a programmer to enable GPU acceleration in the DAGuE framework. The performance achieved by this unchanged code, on a variety of heterogeneous and distributed many cores and GPU resources, demonstrates the desired performance portability.