Radiative Heat Transfer Calculation on 16384 GPUs Using a Reverse Monte Carlo Ray Tracing Approach with Adaptive Mesh Refinement

Radiative Heat Transfer Calculation on 16384 GPUs Using a Reverse Monte Carlo Ray Tracing Approach with Adaptive Mesh Refinement
复制标题

使用具有自适应网格细化的反向蒙特卡洛射线追踪方法在 16384 GPU 上进行辐射传热计算

DOI:
10.1109/ipdpsw.2016.93
复制
发表时间:
2016
期刊:
2016 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
影响因子:
--
通讯作者:
M. Berzins
M. Berzins
中科院分区:
--
文献类型:
--
作者:
A. Humphrey;Daniel Sunderland;T. Harman;M. Berzins

文献摘要

被引文献

相似文献

由于其全对全的物理和计算连接性,模拟热辐射在并行计算方面具有挑战性,并且也是实际应用中的主要传热模式,例如下一代清洁煤锅炉,由Uintah框架建模。然而,在大型计算机系统上,无论是同质的还是异质的,直接的所有对所有的辐射治疗都是非常昂贵的。DOE Titan和计划中的DOE Summit和Sierra机器是当前和新兴的基于GPU的异构系统的例子,其中GPU的处理能力超过CPU的处理能力加剧了这个问题。这些系统要求像Uintah这样的计算框架利用任意数量的节点上GPU,同时在单个模拟中利用数千个GPU。我们表明,辐射热传递问题可以在Uintah内的异构系统通过结合反向蒙特卡罗射线跟踪(RMCRT)技术与AMR的组合,以减少全球通信量。特别是,重要的Uintah基础设施变化,包括一个新的锁和无竞争,线程可扩展的数据结构,用于管理MPI通信请求和改进的内存分配策略,是必要的,以实现出色的强大的扩展结果,以16384 GPU的泰坦。
Modeling thermal radiation is computationally challenging in parallel due to its all-to-all physical and resulting computational connectivity, and is also the dominant mode of heat transfer in practical applications such as next-generation clean coal boilers, being modeled by the Uintah framework. However, a direct all-to-all treatment of radiation is prohibitively expensive on large computers systems whether homogeneous or heterogeneous. DOE Titan and the planned DOE Summit and Sierra machines are examples of current and emerging GPU-based heterogeneous systems where the increased processing capability of GPUs over CPUs exacerbates this problem. These systems require that computational frameworks like Uintah leverage an arbitrary number of on-node GPUs, while simultaneously utilizing thousands of GPUs within a single simulation. We show that radiative heat transfer problems can be made to scale within Uintah on heterogeneous systems through a combination of reverse Monte Carlo ray tracing (RMCRT) techniques combined with AMR, to reduce the amount of global communication. In particular, significant Uintah infrastructure changes, including a novel lock and contention-free, thread-scalable data structure for managing MPI communication requests and improved memory allocation strategies were necessary to achieve excellent strong scaling results to 16384 GPUs on Titan.