Performance Portability of a GPU Enabled Factorization with the DAGuE Framework
Performance Portability of a GPU Enabled Factorization with the DAGuE Framework
复制标题
使用 DAGuE 框架支持 GPU 分解的性能可移植性
DOI:
10.1109/cluster.2011.51
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Jack J. Dongarra
中科院分区:
文献类型:
--
作者:
G. Bosilca;A. Bouteiller;T. Hérault;P. Lemarinier;Narapat Ohm Saengpatsa;S. Tomov;Jack J. Dongarra
Performance portability is a major challenge faced today by developers on heterogeneous high performance computers, consisting of an interconnect, memory with non-uniform access, many-cores and accelerators like GPUs. Recent studies have successfully demonstrated that dense linear algebra operations can be efficiently handled by runtime systems using a DAG representation. In this work, we present the GPU subsystem of the DAGuE runtime, and assess, on the Cholesky factorization test case, the minimal efforts required by a programmer to enable GPU acceleration in the DAGuE framework. The performance achieved by this unchanged code, on a variety of heterogeneous and distributed many cores and GPU resources, demonstrates the desired performance portability.