High performance computing for deformable image registration:: Towards a new paradigm in adaptive radiotherapy

High performance computing for deformable image registration:: Towards a new paradigm in adaptive radiotherapy
复制标题

DOI:
10.1118/1.2948318
复制
发表时间:
2008-08-01
期刊:
影响因子:
3.8
通讯作者:
Owens, John D.
Owens, John D.
中科院分区:
医学3区
文献类型:
--
作者:
Samant, Sanjiv S.;Xia, Junyi;Owens, John D.

文献摘要

被引文献

相似文献

在许多放射治疗中心,随处可得的时间序列成像或时间序列体积成像的出现已经成为治疗计划和适应性放射治疗(ART)不可或缺的组成部分。可变形图像配准(DIR)也被用于医学成像的其他领域,包括运动校正图像重建。由于计算时间较长,DIR在放射治疗和其他领域的临床应用一直受到限制,因此只能进行离线分析。随着硬件和软件的发展,基于图形处理单元(GPU)的计算是一种新兴的通用计算技术,包括DIR,适合于高度并行化的计算。然而,由于受现有编程平台的限制,传统的通用计算在GPU上是有限的。此外,与CPU编程相比,GPU目前的专用处理器内存较少,这可能会限制用于并行处理的有用工作数据集。我们使用NVIDIA 8800 GTX图形处理器和新的CUDA编程语言实现了DEMONS算法。GPU性能将与使用C编程语言的英特尔双核2.4 GHz CPU上的单线程和多线程CPU实现进行比较。CUDA提供了类似C语言的编程接口,并允许直接访问GPU中高度并行的计算单元。并与4DCT获得的临床肺体积图像进行比较。对于图像大小从2.0×10(6)到14.2×10(6)像素的图形处理器,S在1.8-13.5的范围内观察到100次迭代的计算时间。对于单线程实现,GPU注册速度比CPU快55-61倍,对于多线程实现,GPU注册速度快34-39倍。对于基于CPU的计算,计算时间通常与医学成像数据的图像大小成线性关系。计算效率以每百万像素每迭代时间(TPMI)为特征,单位为每百万像素秒(或SPMI)。对于DEMONS算法,我们的CPU实现产生了TPMI的很大程度上的不变值。对于单线程和多线程的情况,平均TPMI分别为0.527 spmi和0.335 spmi,在所考虑的图像数据范围内有2%的变化。对于图形处理器计算,我们得到了tpmi=0.00916 spmi,偏差为3.7%,这表明在CUDA下优化了内存处理。基于GPU的实时DIR的范例为医学成像开辟了一系列临床应用。(C)2008年美国医学物理学家协会。
The advent of readily available temporal imaging or time series volumetric (4D) imaging has become an indispensable component of treatment planning and adaptive radiotherapy (ART) at many radiotherapy centers. Deformable image registration (DIR) is also used in other areas of medical imaging, including motion corrected image reconstruction. Due to long computation time, clinical applications of DIR in radiation therapy and elsewhere have been limited and consequently relegated to offline analysis. With the recent advances in hardware and software, graphics processing unit (GPU) based computing is an emerging technology for general purpose computation, including DIR, and is suitable for highly parallelized computing. However, traditional general purpose computation on the GPU is limited because the constraints of the available programming platforms. As well, compared to CPU programming, the GPU currently has reduced dedicated processor memory, which can limit the useful working data set for parallelized processing. We present an implementation of the demons algorithm using the NVIDIA 8800 GTX GPU and the new CUDA programming language. The GPU performance will be compared with single threading and multithreading CPU implementations on an Intel dual core 2.4 GHz CPU using the C programming language. CUDA provides a C-like language programming interface, and allows for direct access to the highly parallel compute units in the GPU. Comparisons for volumetric clinical lung images acquired using 4DCT were carried out. Computation time for 100 iterations in the range of 1.8-13.5 s was observed for the GPU with image size ranging from 2.0x10(6) to 14.2x10(6) pixels. The GPU registration was 55-61 times faster than the CPU for the single threading implementation, and 34-39 times faster for the multithreading implementation. For CPU based computing, the computational time generally has a linear dependence on image size for medical imaging data. Computational efficiency is characterized in terms of time per megapixels per iteration (TPMI) with units of seconds per megapixels per iteration (or spmi). For the demons algorithm, our CPU implementation yielded largely invariant values of TPMI. The mean TPMIs were 0.527 spmi and 0.335 spmi for the single threading and multithreading cases, respectively, with < 2% variation over the considered image data range. For GPU computing, we achieved TPMI=0.00916 spmi with 3.7% variation, indicating optimized memory handling under CUDA. The paradigm of GPU based real-time DIR opens up a host of clinical applications for medical imaging. (c) 2008 American Association of Physicists in Medicine.