Software technologies coping with memory hierarchy of GPGPU clusters for stencil computations

Software technologies coping with memory hierarchy of GPGPU clusters for stencil computations
复制标题

处理 GPGPU 集群内存层次结构以进行模板计算的软件技术

DOI:
--
复制
发表时间:
2014
期刊:
IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
Guanghao Jin
Guanghao Jin
中科院分区:
--
文献类型:
--
作者:
Toshio Endo;Guanghao Jin

文献摘要

参考文献

被引文献

相似文献

作为CFD模拟的重要核心,网格计算在GPGPU集群上取得了很大的成功,这是由于GPU加速器的高内存带宽和计算速度。然而,计算域的大小受到GPU设备存储器的小容量的限制。为了支持更大的域大小,我们利用GPGPU集群的内存层次结构;更大的主机内存用于维护大域。然而,它是具有挑战性的实现所有更大的域大小,高性能和易于程序开发。为了实现这一目标,我们结合了联合收割机两种软件技术。在算法方面,我们采用了一种局部性改进技术,称为时间分块。在系统软件方面,我们开发了一个MPI/CUDA封装库HHRT,它支持内存交换和细粒度编程模型。有了这种组合,我们证明,我们的目标是通过TSUBAME2.5,千万亿次GPGPU超级计算机上的评估实现。
Stencil computations, which are important kernels for CFD simulations, have been highly successful on GPGPU clusters, due to high memory bandwidth and computation speed of GPU accelerators. However, sizes of the computed domains are limited by small capacity of GPU device memory. In order to support larger domain sizes, we utilize the memory hierarchy of GPGPU clusters; larger host memory is used for maintain large domains. However, it is challenging to achieve all of larger domain sizes, high performance and easiness of program development. Towards this goal, we combine two software technologies. From the aspect of algorithm, we adopt a locality improvement technique called temporal blocking. From the aspect of system software, we developed a MPI/CUDA wrapper library named HHRT, which supports memory swapping and finer grained programming model. With this combination, we demonstrate that our goal is achieved through evaluations on TSUBAME2.5, a petascale GPGPU supercomputer.
用于并行计算的语言和编译器
DOI: 10.1007/978-3-540-85261-2_12
发表时间: 2008
期刊: --
影响因子: --
作者:
Cornwall J
通讯作者: Cornwall J