Systematic Performance Optimization of Cone-Beam Back-Projection on the Kepler Architecture

Systematic Performance Optimization of Cone-Beam Back-Projection on the Kepler Architecture
复制标题

开普勒架构上锥束反投影的系统性能优化

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
B. Keck
B. Keck
中科院分区:
--
文献类型:
--
作者:
T. Zinßer;B. Keck

文献摘要

被引文献

相似文献

滤波反投影算法广泛用于从介入 C 形臂计算机断层扫描中的锥束投影重建体积数据。此外,通用 GPU 已成为在时间紧迫的临床过程中加速重建的流行工具。在这项工作中,我们专注于在支持 CUDA 的 GPU 的最新架构上对锥束背投影进行系统性能优化。我们的优化方法基于通过分析专门修改的内核来识别主要性能瓶颈。我们的主要贡献是对反投影算法的智能重构,有利于同时处理大量投影,同时提高纹理缓存的命中率。我们使用著名的 RabbitCT 基准测试来展示我们在单个基于 Kepler 的 GeForce GTX 680 GPU 上实现的出色性能。我们的实现在不到一秒的时间内将 496 个输入投影反向投影到 5123 立方体积上,这是最佳竞争实现的三倍。我们的反投影实现还能够在大约六秒内重建立方体 10243 体积,这是我们已知的最佳竞争实现的六倍。
Filtered back-projection algorithms are widely used for the reconstruction of volumetric data from cone-beam projections in interventional C-arm computed tomography. Furthermore, general-purpose GPUs have become a popular tool for accelerating the reconstruction during time-critical clinical procedures. In this work, we focus on the systematic performance optimization of cone-beam back-projection on the latest architecture of CUDA-enabled GPUs. Our optimization approach is based on the identification of the major performance bottleneck through the analysis of specifically modified kernels. Our main contribution is a smart restructuring of the backprojection algorithm that facilitates the simultaneous processing of a large number of projections and improves the hit rate of the texture cache at the same time. We use the well-known RabbitCT benchmark to demonstrate the outstanding performance of our implementation on a single Kepler-based GeForce GTX 680 GPU. Our implementation performs the back-projection of 496 input projections onto a cubic 5123 volume in less than one second, which is three times as fast as the best competing implementation. Our back-projection implementation is also able to reconstruct a cubic 10243 volume in about six seconds, which is six times as fast as the best competing implementation known to us.