Decomposition method for fast computation of gigapixel-sized Fresnel holograms on a graphics processing unit cluster

Decomposition method for fast computation of gigapixel-sized Fresnel holograms on a graphics processing unit cluster
复制标题

DOI:
10.1364/ao.57.003134
复制
发表时间:
2018-04-20
期刊:
影响因子:
1.9
通讯作者:
Baba, Takanobu
Baba, Takanobu
中科院分区:
工程技术4区
文献类型:
--
作者:
Jackin, Boaz Jessie;Watanabe, Shinpei;Baba, Takanobu

文献摘要

被引文献

相似文献

提出了一种大尺寸菲涅耳计算全息图的并行计算方法。该方法介绍了我们在早期的报告作为一种技术计算傅立叶计算全息图从二维物体数据。本文将该方法推广到由三维物体数据计算菲涅耳计算全息图。计算问题的规模也扩大到2千兆像素,使其更接近真实的应用需求。该方法的显著特点是能够避免通信开销,从而充分利用并行设备的计算能力。该方法具有三层并行性,有利于小型到大型并行计算机。仿真和光学实验进行了演示的可行性和评估所提出的技术的效率。与传统方法相比,在16节点集群(每个节点一个GPU)上仅利用一层并行性,计算速度提高了两倍。一个20倍的提高计算速度已估计利用两层的并行在一个非常大规模的并行机与16个节点,其中每个节点有16个GPU。(C)2018美国光学学会
A parallel computation method for large-size Fresnel computer-generated hologram (CGH) is reported. The method was introduced by us in an earlier report as a technique for calculating Fourier CGH from 2D object data. In this paper we extend the method to compute Fresnel CGH from 3D object data. The scale of the computation problem is also expanded to 2 gigapixels, making it closer to real application requirements. The significant feature of the reported method is its ability to avoid communication overhead and thereby fully utilize the computing power of parallel devices. The method exhibits three layers of parallelism that favor small to large scale parallel computing machines. Simulation and optical experiments were conducted to demonstrate the workability and to evaluate the efficiency of the proposed technique. A two-times improvement in computation speed has been achieved compared to the conventional method, on a 16-node cluster (one GPU per node) utilizing only one layer of parallelism. A 20-times improvement in computation speed has been estimated utilizing two layers of parallelism on a very large-scale parallel machine with 16 nodes, where each node has 16 GPUs. (C) 2018 Optical Society of America.