Fast and Parallel Computation of the Discrete Periodic Radon Transform on GPUs, Multicore CPUs and FPGAs

Fast and Parallel Computation of the Discrete Periodic Radon Transform on GPUs, Multicore CPUs and FPGAs
复制标题

GPU、多核 CPU 和 FPGA 上离散周期 Radon 变换的快速并行计算

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Information Photonics
影响因子:
--
通讯作者:
D. Llamocca
D. Llamocca
中科院分区:
--
文献类型:
--
作者:
Cesar Carranza;M. Pattichis;D. Llamocca

文献摘要

被引文献

相似文献

离散周期Radon变换(DPRT)在从投影重建图像中有许多重要的应用,最近已被用于计算2D卷积的快速和可扩展的架构中。不幸的是,DPRT的直接计算涉及$O(N^{3})$加法和存储器访问,这在单核架构中可能非常昂贵。本文提出了在多核CPU和GPU上计算DPRT及其逆的新的高效算法。将结果与专用硬件实现(FPGA/ASIC)进行比较。这些结果为新算法的成功提供了重要证据。在8核CPU(Intel Xeon)上,支持每核两个线程,FastDirDPRT和FastDirInvDPRT实现了约10美元的加速 imes(mathbf{up} mathbf{to} 12.83 imes)$的性能。在2048核GPU(GTX 980)上,FastRayDPRT和FastRayInvDPRT实现了526范围内的加速比(售价127美元 imes 127$)至873(1021美元 IMES 1021$),其近似于可以实现的理想加速比。DPRT可以精确地实时计算(每秒30帧),价格为1471美元 imes 1471$图像使用FastRayDPRT上的GPU。此外,GPU算法近似于使用$2N$并行内核在100 MHz下的高效FPGA实现的性能。
The Discrete Periodic Radon Transform (DPRT) has many important applications in reconstructing images from their projections and has recently been used in fast and scalable architectures for computing 2D convolutions. Unfortunately, the direct computation of the DPRT involves $O(N^{3})$ additions and memory accesses that can be very costly in single-core architectures. The current paper presents new and efficient algorithms for computing the DPRT and its inverse on multi-core CPUs and GPUs. The results are compared against specialized hardware implementations (FPGAs/ASICs). The results provide significant evidence of the success of the new algorithms. On an 8-core CPU (Intel Xeon), with support for two threads per core, FastDirDPRT and FastDirInvDPRT achieve a speedup of approximately $10 imes (mathbf{up} mathbf{to} 12.83 imes)$ over the single-core CPU implementation. On a 2048-core GPU (GTX 980), FastRayDPRT and FastRayInvDPRT achieve speedups in the range of 526 (for $127 imes 127$) to 873 (for $1021 imes 1021$), which approximate ideal speedups of what can be achieved. The DPRT can be computed exactly and in real-time (30 frames per second) for $1471 imes 1471$ images using FastRayDPRT on the GPU. Furthermore, the GPU algorithms approximate the performance of an efficient FPGA implementation using $2N$ parallel cores at 100MHz.