High-performance tomographic reconstruction using graphics processing units

High-performance tomographic reconstruction using graphics processing units
复制标题

使用图形处理单元进行高性能断层扫描重建

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
T. Gureyev
T. Gureyev
中科院分区:
--
文献类型:
--
作者:
Y. Nesterets;T. Gureyev

文献摘要

被引文献

相似文献

计算机断层扫描 (CT) 已成为对可见光不透明的物体内部结构进行三维 (3D) 可视化的常规工具。 X 射线 CT 广泛应用于材料科学、生物医学应用和其他领域。人们已经开发了不同的 CT 重建算法,包括平行束几何中著名的滤波反投影 (FBP) 算法和适用于具有小型 X 射线源和平面二维 (2D) 探测器的圆形轨迹的锥束几何的 Feldkamp-Davis-Kress (FDK) 算法。探测器技术的最新进展使得线性尺寸达到 4k 像素或更高的商用 2D 电荷耦合器件 (CCD) 得以面世。存储如此线性大小的 3D 浮点数据量所需的内存量为 256GB 或更多,这大大超过了高端台式计算机和小型计算机集群中常见的 RAM 量。 FBP 和 FDK CT 重建算法中计算量最大的步骤是所谓的反投影操作,在典型的基于 CPU 的实现中,该操作占用了总重建时间的 99%。通用图形处理单元 (GP-GPU) 的 3D 计算机图形功能例如在过去的十到十五年里,通过 OpenGL 或 DirectX 一直用于背投影操作。随着最近重建体积大小的增加,如上所述,通常将整个重建体积存储在 GPU 内存中的标准方法已经成为问题。有效使用 GPU 进行大数据量 CT 重建需要不同的算法方法。我们开发了新的基于 CPU 和 GPU 的 X 射线 CT FBP 和 FDK 算法实现。这些实现考虑了以下原则: • 使用尽可能少的 RAM 和/或 GPU 内存来重建对象的每个轴向切片。例如,对于基于 GPU 的反投影代码,在 GPU 上仅分配用于单个重构切片的内存;这允许使用高端 GPU(板载超过 1GB 内存)重建线性尺寸高达 16k 的卷; • 每个轴向切片的重建应尽可能独立于其他切片。这允许使用 CPU/GPU 多线程功能并行重建不同的切片。因此,总重建时间随着 CPU 核心和/或 GPU 数量的增加而减少。根据我们的测试,与单个 CPU 核的相应结果相比,基于 GPU 的反投影操作实现可将反投影本身加速两个数量级,并将总 CT 重建加速一个数量级以上。
Computed Tomography (CT) has become a routine tool for three-dimensional (3D) visualization of the internal structure of objects which are opaque to visible light. X-ray CT is widely used in materials science, biomedical applications and elsewhere. Different CT reconstruction algorithms have been developed including the well-known Filtered-Back-Projection (FBP) algorithm in the parallel beam geometry and the Feldkamp-Davis-Kress (FDK) algorithm applicable to cone-beam geometry with a circular trajectory of a small X-ray source and a flat two-dimensional (2D) detector. Recent progress in detector technology resulted in the availability of commercial 2D Charge-Coupled-Devices (CCDs) with linear dimensions of the order of 4k pixels or more. The amount of memory required for storing a 3D volume of floating-point data of such linear size is 256GB or more, which significantly exceeds the typical amount of RAM found not only in high-end desktop computers but also in small computer clusters. The most computationally intensive step in the FBP and FDK CT reconstruction algorithms is the so-called back-projection operation that takes up to 99% of the total reconstruction time in a typical CPU-based implementation. The 3D computer graphics capabilities of general-purpose graphics processing units (GP- GPUs) utilized e.g. via OpenGL or DirectX have been used for the back-projection operation for the last ten to fifteen years. With the recent increase in the size of reconstructed volumes as mentioned above, standard approaches which usually store the whole reconstructed volume in GPU memory have become problematic. Different algorithmic approaches are required for the effective use of GPUs for CT reconstruction of large data volumes. We have developed new CPU-based and GPU-based implementations of the FBP and FDK algorithms for X- ray CT. These implementations take into account the following principles: • Use as little RAM and/or GPU memory for the reconstruction of each axial slice of the object as possible. E.g., for the GPU-based back-projection code, memory for only a single reconstructed slice is allocated on the GPU; this allows the reconstruction of volumes with linear dimension of up to 16k using top-end GPUs (with more than 1GB memory onboard); • Reconstruction of each axial slice should be as independent from the others as possible. This allows for parallel reconstruction of different slices using CPU/GPU multithreading capabilities. As a result, the total reconstruction time reduces with the number of CPU cores and/or GPUs. According to our tests, the GPU-based implementations of the back-projection operation result in up to two orders of magnitude speed-up of the back-projection itself and more than an order of magnitude speed-up of the total CT reconstruction compared to the corresponding results for a single CPU core.