High-performance cone beam reconstruction using CUDA compatible GPUs

High-performance cone beam reconstruction using CUDA compatible GPUs
复制标题

DOI:
10.1016/j.parco.2010.01.004
复制
发表时间:
2010-02-01
期刊:
影响因子:
1.4
通讯作者:
Hagihara, Kenichi
Hagihara, Kenichi
中科院分区:
计算机科学4区
文献类型:
--
作者:
Okitsu, Yusuke;Ino, Fumihiko;Hagihara, Kenichi

文献摘要

被引文献

相似文献

计算统一设备架构(CUDA)是一个软件开发平台,允许我们在nVIDIA图形处理单元(CPU)上运行类似C的程序。提出了一种基于CUDA兼容GPU的锥束重建加速方法。所提出的方法使用三种技术加速Feldkamp,Davis和Kress(FDK)算法:(1)片外存储器访问减少以节省存储器带宽;(2)循环展开以隐藏存储器延迟;以及(3)多线程以利用多个GPU。我们描述了如何将这些技术可以被纳入重建代码。我们还展示了一个分析模型,以了解多GPU环境下的重建性能。实验结果表明,该方法运行在83%的理论内存带宽,实现了64.3投影每秒(pps)的吞吐量为512(3)-体素体积从360 512(2)-像素的投影重建。这种性能比以前基于CUDA的方法高出41%,比基于CPU的方法快24倍,该方法通过向量内部函数进行优化。还提出了一些详细的分析,以了解如何有效地加速技术提高重建性能的一个天真的方法。我们还展示了大规模数据集的核心外重建,高达1024(3)体素体积。(C)2010 Elsevier B. V.保留所有权利。
Compute unified device architecture (CUDA) is a software development platform that allows us to run C-like programs on the nVIDIA graphics processing unit (CPU). This paper presents an acceleration method for cone beam reconstruction using CUDA compatible GPUs. The proposed method accelerates the Feldkamp, Davis, and Kress (FDK) algorithm using three techniques: (1) off-chip memory access reduction for saving the memory bandwidth; (2) loop unrolling for hiding the memory latency; and (3) multithreading for exploiting multiple GPUs. We describe how these techniques can be incorporated into the reconstruction code. We also show an analytical model to understand the reconstruction performance on multi-GPU environments. Experimental results show that the proposed method runs at 83% of the theoretical memory bandwidth, achieving a throughput of 64.3 projections per second (pps) for reconstruction of 512(3)-voxel volume from 360 512(2)-pixel projections. This performance is 41% higher than the previous CUDA-based method and is 24 times faster than a CPU-based method optimized by vector intrinsics. Some detailed analyses are also presented to understand how effectively the acceleration techniques increase the reconstruction performance of a naive method. We also demonstrate out-of-core reconstruction for large-scale datasets, up to 1024(3)-voxel volume. (C) 2010 Elsevier B.V. All rights reserved.