Heterogeneous concurrent execution of Monte Carlo photon transport on CPU, GPU and MIC

Heterogeneous concurrent execution of Monte Carlo photon transport on CPU, GPU and MIC
复制标题

CPU、GPU 和 MIC 上蒙特卡罗光子传输的异构并发执行

DOI:
10.1109/ia3.2014.11
复制
发表时间:
2014
期刊:
2014 23rd International Conference on Parallel Architecture and Compilation (PACT)
影响因子:
--
通讯作者:
X. Xu
X. Xu
中科院分区:
--
文献类型:
--
作者:
Noah Wolfe;Tianyu Liu;C. Carothers;X. Xu

文献摘要

被引文献

相似文献

本文提出了一种新的蒙特卡罗光子传输的异构型并发执行方法。Archer,一个用于计算全身患者模体CT成像辐射剂量的应用程序已经扩展到在CPU、GPU和MIC的任何组合上同时执行。Archer的目标是检测并同时利用所有可用的CPU、GPU和MIC处理设备。由于蒙特卡罗光子传输算法的不规则性,我们实现了一种新的“自助式”方法来组织异质设备计算。这种方法高效且有效地允许每个设备重复抓取域的部分并并行计算,直到整个域被模拟。使用各种Intel和NVIDIA设备的各种组合制作并展示了新的时序基准测试。与仅使用Intel Xeon X5650相比,同时使用Intel Xeon X5650 CPU、Intel Xeon Phi 5110P MIC和NVIDIA K40 GPU的速度提高了13倍。
In this paper, a new level of heterogeneous concurrent execution of Monte Carlo photon transport is presented. ARCHER, an application for computing radiation dosimetry for CT imaging involving whole-body patient phantoms has been extended to execute on any combination of CPUs, GPUs and MICs concurrently. The goal is for ARCHER to detect and simultaneously utilize all CPU, GPU and MIC processing devices available. Due to the irregular nature of the Monte Carlo photon transport algorithm, a new "self service" approach to organizing the heterogeneous device computing has been implemented. This approach efficiently and effectively allows each device to repeatedly grab portions of the domain and compute concurrently until the entire domain has been simulated. New timing benchmarks using various combinations of various Intel and NVIDIA devices are made and presented. A speedup of 13x has been observed when utilizing Intel's Xeon X5650 CPU, Intel's Xeon Phi 5110P MIC and NVIDIA's K40 GPU concurrently versus just the Intel Xeon X5650.