Efficient methods for implementation of multi-level nonrigid mass-preserving image registration on GPUs and multi-threaded CPUs.

Efficient methods for implementation of multi-level nonrigid mass-preserving image registration on GPUs and multi-threaded CPUs.
复制标题

DOI:
10.1016/j.cmpb.2015.12.018
复制
发表时间:
2016-04
影响因子:
6.1
通讯作者:
Lin CL
Lin CL
中科院分区:
工程技术2区
文献类型:
--
作者:
Ellingwood ND;Yin Y;Smith M;Lin CL

文献摘要

被引文献

相似文献

更快、更准确的图像配准方法对于利用医学成像进行基于人群的研究以及改进临床应用非常重要。我们提出了一种新的计算和内存高效的多层次的方法,图形处理单元(GPU)进行注册的两个计算机断层扫描(CT)体积肺图像。我们开发了一种计算和存储效率高的同构多层B样条变换复合(DMTC)方法,在GPU上实现了两幅CT肺部图像的非刚性质量保持配准。该框架由一个层次的B样条控制网格的分辨率不断增加。一个相似性标准被称为平方组织体积差(SSTVD)的总和被用来保存肺组织质量。SSTVD的使用包括计算组织体积、雅可比矩阵及其导数,由于内存限制,这使得其在GPU上的实现具有挑战性。DMTC方法的使用使得能够减少变量的计算和存储器存储,并且由于预计算值的能力,GPU和中央处理单元(CPU)之间的通信最少。该方法在6名健康受试者身上进行了评估。将所得GPU生成的位移场与先前确认的CPU对应场进行比较,结果显示出良好的一致性,平均归一化均方根误差(nRMS)为0.044 ± 0.015。比较了单线程CPU、多线程CPU和GPU算法的并行性和性能加速比。最佳性能加速发生在GPU实现的SSTVD成本和成本梯度计算的最高分辨率下,当考虑使用Nvidia Tesla K20X GPU的每次迭代的平均时间时,加速是单线程CPU版本的112倍,是12线程版本的11倍。所提出的基于GPU的DMTC方法在运行时间方面优于其多线程CPU版本。在GPU版本上,总注册时间减少到2.9分钟,而在12线程CPU版本上为12.8分钟,在单线程CPU上为112.5分钟。此外,本工作中讨论的GPU实现可以适用于需要计算一阶导数的其他成本函数。
Faster and more accurate methods for registration of images are important for research involved in conducting population-based studies that utilize medical imaging, as well as improvements for use in clinical applications. We present a novel computation- and memory-efficient multi-level method on graphics processing units (GPU) for performing registration of two computed tomography (CT) volumetric lung images. We developed a computation- and memory-efficient Diffeomorphic Multi-level B-Spline Transform Composite (DMTC) method to implement nonrigid mass-preserving registration of two CT lung images on GPU. The framework consists of a hierarchy of B-Spline control grids of increasing resolution. A similarity criterion known as the sum of squared tissue volume difference (SSTVD) was adopted to preserve lung tissue mass. The use of SSTVD consists of the calculation of the tissue volume, the Jacobian, and their derivatives, which makes its implementation on GPU challenging due to memory constraints. The use of the DMTC method enabled reduced computation and memory storage of variables with minimal communication between GPU and Central Processing Unit (CPU) due to ability to pre-compute values. The method was assessed on six healthy human subjects. Resultant GPU-generated displacement fields were compared against the previously validated CPU counterpart fields, showing good agreement with an average normalized root mean square error (nRMS) of 0.044 ± 0.015. Runtime and performance speedup are compared between single-threaded CPU, multi-threaded CPU, and GPU algorithms. Best performance speedup occurs at the highest resolution in the GPU implementation for the SSTVD cost and cost gradient computations, with a speedup of 112 times that of the single-threaded CPU version and 11 times over the twelve-threaded version when considering average time per iteration using a Nvidia Tesla K20X GPU. The proposed GPU-based DMTC method outperforms its multi-threaded CPU version in terms of runtime. Total registration time reduced runtime to 2.9 min on the GPU version, compared to 12.8 min on twelve-threaded CPU version and 112.5 min on a single-threaded CPU. Furthermore, the GPU implementation discussed in this work can be adapted for use of other cost functions that require calculation of the first derivatives.