Exploiting Hardware-Accelerated Ray Tracing for Monte Carlo Particle Transport with OpenMC

Exploiting Hardware-Accelerated Ray Tracing for Monte Carlo Particle Transport with OpenMC
复制标题

利用 OpenMC 硬件加速光线追踪进行蒙特卡罗粒子传输

DOI:
10.1109/pmbs49563.2019.00008
复制
发表时间:
2019
期刊:
2019 IEEE/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)
影响因子:
--
通讯作者:
Simon McIntosh
Simon McIntosh
中科院分区:
--
文献类型:
--
作者:
Justin Salmon;Simon McIntosh

文献摘要

被引文献

相似文献

OpenMC是一种基于CPU的Monte Carlo粒子传输模拟代码,最近在MIT的计算机反应堆物理组中开发了,目前正在英国原子能授权授权授权使用ITER FUSION反应器项目。在NVIDIA的最新图灵GPU中展示了一个新的OpenMC港口,以在新的Ray Tracing(RT)核心上运行。 9.8倍在16核CPU上使用天然建设性固体几何形状,并使用近似三角形网状几何形状进行13倍的速度。一个射线追踪问题,有机会通过利用Turing GPU中的RT核心来启用硬件加速射线跟踪,从而在三角形网眼上获得更高的性能。扩展GPU端口以支持RT核心加速度在2倍和20倍之间,我们注意到几何模型的复杂性对性能有重大影响,因为RT核心加速度会随着我们所知名度的增加而产生的速度相对较高。是第一个工作表明,我们可以根据Wilder得出有关RT核心的科学工作负载的剥削。适用性,限制和性能可移植性。
OpenMC is a CPU-based Monte Carlo particle transport simulation code recently developed in the Computa- tional Reactor Physics Group at MIT, and which is currently being evaluated by the UK Atomic Energy Authority for use on the ITER fusion reactor project. In this paper we present a novel port of OpenMC to run on the new ray tracing (RT) cores in NVIDIA’s latest Turing GPUs. We show here that the OpenMC GPU port yields up to 9.8x speedup on a single node over a 16-core CPU using the native constructive solid geometry, and up to 13x speedup using approximate triangle mesh geometry. Furthermore, since the expensive 3D geometric operations re- quired during particle transport simulation can be formulated as a ray tracing problem, there is an opportunity to gain even higher performance on triangle meshes by exploiting the RT cores in Turing GPUs to enable hardware-accelerated ray tracing. Extending the GPU port to support RT core acceleration yields between 2x and 20x additional speedup. We note that geometric model complexity has a significant impact on performance, with RT core acceleration yielding comparatively greater speedups as complexity increases. To the best of our knowledge, this is the first work showing that exploitation of RT cores for scientific workloads is possible. We finish by drawing conclusions about RT cores in terms of wider applicability, limitations and performance portability.