Next-generation acceleration and code optimization for light transport in turbid media using GPUs.

Next-generation acceleration and code optimization for light transport in turbid media using GPUs.
复制标题

DOI:
10.1364/boe.1.000658
复制
发表时间:
2010-09-01
影响因子:
3.4
通讯作者:
Lilge L
Lilge L
中科院分区:
医学2区
文献类型:
--
作者:
Alerstam E;Lo WC;Han TD;Rose J;Andersson-Engels S;Lilge L

文献摘要

被引文献

相似文献

一个高度优化的蒙特卡罗(MC)程序包,用于模拟光 传输是在最新的图形处理单元(GPU)上开发的, 用于NVIDIA的通用计算-Fermi GPU。在生物医学 光学,MC方法是模拟光的黄金标准方法 生物组织中的运输,由于其准确性和灵活性 在3D中模拟逼真的异质组织几何形状。但 MC模拟在反问题中的广泛使用,例如治疗 PDT的规划受到计算时间长的限制。尽管 并行性,在GPU上优化MC代码已被证明是一种 挑战,特别是当共享模拟结果矩阵之间 许多并行线程需要频繁使用原子指令, 访问慢速GPU全局内存。本文提出了一种优化 一种利用快速共享内存解决性能问题的方案 原子访问造成的瓶颈,并讨论了许多其他 充分利用GPU潜力所需的优化技术。 使用这些技术,生物光子学中广泛接受的MC代码包, 称为MCML,在Fermi GPU上成功加速了大约 与最先进的英特尔酷睿i7 CPU相比,性能提升了600倍。皮肤模型 由7层组成的层用作标准模拟几何形状。到 演示GPU集群计算的可能性,相同的GPU代码 在四个GPU上执行,显示出性能的线性改善, 越来越多的GPU。基于GPU的MCML代码包,名为GPU-MCML, 与各种图形卡兼容,并作为 开源软件有两个版本:一个优化的版本, 性能和简化版本的初学者()。
A highly optimized Monte Carlo (MC) code package for simulating light transport is developed on the latest graphics processing unit (GPU) built for general-purpose computing from NVIDIA - the Fermi GPU. In biomedical optics, the MC method is the gold standard approach for simulating light transport in biological tissue, both due to its accuracy and its flexibility in modelling realistic, heterogeneous tissue geometry in 3-D. However, the widespread use of MC simulations in inverse problems, such as treatment planning for PDT, is limited by their long computation time. Despite its parallel nature, optimizing MC code on the GPU has been shown to be a challenge, particularly when the sharing of simulation result matrices among many parallel threads demands the frequent use of atomic instructions to access the slow GPU global memory. This paper proposes an optimization scheme that utilizes the fast shared memory to resolve the performance bottleneck caused by atomic access, and discusses numerous other optimization techniques needed to harness the full potential of the GPU. Using these techniques, a widely accepted MC code package in biophotonics, called MCML, was successfully accelerated on a Fermi GPU by approximately 600x compared to a state-of-the-art Intel Core i7 CPU. A skin model consisting of 7 layers was used as the standard simulation geometry. To demonstrate the possibility of GPU cluster computing, the same GPU code was executed on four GPUs, showing a linear improvement in performance with an increasing number of GPUs. The GPU-based MCML code package, named GPU-MCML, is compatible with a wide range of graphics cards and is released as an open-source software in two versions: an optimized version tuned for high performance and a simplified version for beginners ().