Data Distribution Method for Fast Giga-scale Hologram Generation on a Multi-GPU Cluster

Data Distribution Method for Fast Giga-scale Hologram Generation on a Multi-GPU Cluster
复制标题

DOI:
10.1145/3231104.3231105
复制
发表时间:
2018-07
期刊:
Proceedings of the 2018 Workshop on Advanced Tools, Programming Languages, and PLatforms for Implementing and Evaluating Algorithms for Distributed systems
影响因子:
--
通讯作者:
T. Baba;Shinpei Watanabe;B. Jackin;K. Ootsu;Takeshi Ohkawa;T. Yokota;Y. Hayasaki;T. Yatagai
T. Baba;Shinpei Watanabe;B. Jackin;K. Ootsu;Takeshi Ohkawa;T. Yokota;Y. Hayasaki;T. Yatagai
中科院分区:
其他
文献类型:
--
作者:
T. Baba;Shinpei Watanabe;B. Jackin;K. Ootsu;Takeshi Ohkawa;T. Yokota;Y. Hayasaki;T. Yatagai

文献摘要

相似文献

3D全息显示器长期以来一直被视为未来的人机界面,因为它不需要用户佩戴特殊设备。然而,除了显示设备技术的延迟之外,其繁重的计算要求阻碍了这种显示的实现。最近的一项研究表明,应该实时处理具有数十亿像素的物体和全息图,以实现高分辨率和宽视角。针对这个问题,首先,我们提出了一种新的数据分发方法,该方法利用基于 FFT 的基本 O(N log N) 计算,但在多 GPU 集群上的计算过程中不需要任何节点间通信。然后,我们在多 GPU 集群上实现了该方法,应用了多种单节点和多节点优化和并行化技术。实验结果表明,节点内优化比原始单节点代码获得了11.52倍的加速。此外,使用 8 个节点、每个节点 2 个 GPU 的多节点优化实现了 4.28 秒的执行时间。用于从 3.2 GB 像素物体生成 1.6 GB 像素全息图。这意味着CPU使用传统的基于FFT的算法的顺序处理速度提高了237.92倍。
The 3D holographic display has long been expected as a future human interface as it does not require users to wear special devices. However, in addition to the delay of display device technology, its heavy computation requirement prevents the realization of such displays. A recent study says that objects and holograms with several giga-pixels should be processed in real time for the realization of high resolution and wide view angle. To this problem, first, we have proposed a new data distribution method that utilizes a basic FFT-based O(N log N) computation but does not need any inter-node communications during the computation on a multi-GPU cluster. Then, we have implemented the method on a multi-GPU cluster, applying several single-node and multi-node optimization and parallelization techniques. The experimental results show that the intra-node optimizations attain 11.52 times speed-up from the original single node code. Further, multi-node optimizations using 8 nodes, 2 GPUs per node, attain the execution time of 4.28 sec. for generating 1.6 giga-pixel hologram from 3.2 giga-pixel object. It means 237.92 times speed-up of the sequential processing by CPU using a conventional FFT-based algorithm.