Fast Computation with Efficient Object Data Distribution for Large-Scale Hologram Generation on a Multi-GPU Cluster

Fast Computation with Efficient Object Data Distribution for Large-Scale Hologram Generation on a Multi-GPU Cluster
复制标题

DOI:
10.1587/transinf.2018edp7346
复制
发表时间:
2019-07
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
T. Baba;Shinpei Watanabe;B. Jackin;K. Ootsu;Takeshi Ohkawa;T. Yokota;Y. Hayasaki;T. Yatagai
T. Baba;Shinpei Watanabe;B. Jackin;K. Ootsu;Takeshi Ohkawa;T. Yokota;Y. Hayasaki;T. Yatagai
中科院分区:
其他
文献类型:
--
作者:
T. Baba;Shinpei Watanabe;B. Jackin;K. Ootsu;Takeshi Ohkawa;T. Yokota;Y. Hayasaki;T. Yatagai

文献摘要

相似文献

3D全息显示一直被认为是未来的人机界面,因为它不需要用户佩戴特殊设备。然而,其繁重的计算要求阻止了这种显示器的实现。最近的一项研究表明,为了实现高分辨率和宽视角,需要对具有数十亿像素的物体和全息图进行真实的实时处理。针对这个问题,首先,我们采用了传统的FFT算法的GPU集群环境,以避免繁重的节点间通信。然后,我们应用了几种单节点和多节点优化和并行化技术。单节点优化包括改变对象分解方式、减少CPU和GPU之间的数据传输、内核集成、流处理以及在一个节点内利用多个GPU。多节点优化包括对象数据从主机节点到其他节点的分发方法。实验结果表明,节点内优化获得了11.52倍的速度比原来的单节点代码。此外,使用8个节点(每个节点2个GPU)的多节点优化实现了4.28秒的执行时间,用于从3.2千兆像素的对象生成1.6千兆像素的全息图。在传统的FFT算法下,CPU的顺序处理速度提高了237.92倍,多核CPU的多线程执行速度提高了41.78倍。关键词:计算全息,大规模计算全息,GPU集群
The 3D holographic display has long been expected as a future human interface as it does not require users to wear special devices. However, its heavy computation requirement prevents the realization of such displays. A recent study says that objects and holograms with several giga-pixels should be processed in real time for the realization of high resolution and wide view angle. To this problem, first, we have adapted a conventional FFT algorithm to a GPU cluster environment in order to avoid heavy inter-node communications. Then, we have applied several single-node and multi-node optimization and parallelization techniques. The single-node optimizations include a change of the way of object decomposition, reduction of data transfer between the CPU and GPU, kernel integration, stream processing, and utilization of multiple GPUs within a node. The multi-node optimizations include distribution methods of object data from host node to the other nodes. Experimental results show that intra-node optimizations attain 11.52 times speed-up from the original single node code. Further, multi-node optimizations using 8 nodes, 2 GPUs per node, attain an execution time of 4.28 sec for generating a 1.6 giga-pixel hologram from a 3.2 giga-pixel object. It means a 237.92 times speed-up of the sequential processing by CPU and 41.78 times speed-up of multi-threaded execution on multicore-CPU, using a conventional FFTbased algorithm. key words: computer generated holography, large-scale CGH, GPU cluster