Using Multiple GPUs to Accelerate MTF Compensation and Georectification of High-Resolution Optical Satellite Images

Using Multiple GPUs to Accelerate MTF Compensation and Georectification of High-Resolution Optical Satellite Images
复制标题

DOI:
10.1109/jstars.2015.2477460
复制
发表时间:
2015-10
影响因子:
5.5
通讯作者:
Mi Wang;Liuyang Fang;Deren Li;Jun Pan
Mi Wang;Liuyang Fang;Deren Li;Jun Pan
中科院分区:
工程技术3区
文献类型:
--
作者:
Mi Wang;Liuyang Fang;Deren Li;Jun Pan

文献摘要

被引文献

相似文献

现代高分辨率光学卫星收集的数据量迅速增长,给近实时处理带来了压力。在本文中,我们提出了我们最近的工作的调制传递函数补偿(MTFC)和地理校正(GR),两个最耗时的光学卫星图像处理算法,使用多个图形处理单元(多GPU)的加速。试验使用了由10幅ZY-3天底图组成的特制条带,覆盖了台风“菲特”造成的大部分灾区(ZY-3是中国第一颗高精度民用立体测绘光学卫星)。算法的快速分析表明,补偿和校正占用了MTFC和GR总运行时间的99.50%以上。为了缩短时间,我们将这两个操作移植到一个多GPU系统中,该系统由一个英特尔酷睿i7 CPU和三个费米架构NVIDIA GTX 580 GPU组成。首先,在基本单GPU实现的早期阶段确定内核排列和初始设置。二是三项优化措施,采用最大化存储器吞吐量、优化流控制指令以及重叠数据传送和内核执行来进一步提高性能。实验取得了显着的加速比MTFC和GR分别为102.9和184.2。接下来,介绍两种多GPU策略,即,合作处理(CP)和独立处理(IP)。实验结果表明,当处理图像的数量是GPU数量的倍数时,IP是最佳选择;当处理图像的数量是GPU数量的倍数时,CP是最佳选择。此外,Intel Core i7和NVIDIA GTX 580都完全支持IEEE 754-2008浮点精度标准;因此,我们的GPU实现的正确性可以得到充分保证。
The rapid growth in the volume of data collected by modern high-resolution optical satellites puts pressure on near real-time processing. In this paper, we present our recent work on the acceleration of modulation transfer function compensation (MTFC) and georectification (GR), two of the most time-consuming optical satellite image processing algorithms, using multiple graphic processing units (multi-GPUs). A tailored strip consisting of 10 ZY-3 nadir images and covering most of the disaster area caused by Typhoon Fitow is used for the experiment (ZY-3 is the first high-accuracy civilian stereo-mapping optical satellite of China). Rapid profiling of the algorithms reveals that compensation and rectification take virtually over 99.50% of the total run times of MTFC and GR. To shorten the time, we port these two operations to a multi-GPU system that consists of an Intel Core i7 CPU and three Fermi-architecture NVIDIA GTX 580 GPUs. First, kernel arrangement and initial settings are determined in the early stage for basic single-GPU implementation. Second, three optimization measures, i.e., maximizing memory throughput, optimizing flow control instructions, and overlapping data transfer and kernel execution, are taken to further improve performance. The experiments achieved significant speedup ratios of 102.9 and 184.2 for MTFC and GR, respectively. Next, two multi-GPU strategies, i.e., cooperative processing (CP) and independent processing (IP), are proposed. The experimental results show that IP is the best option if the number of images to be processed is a multiple of the number of GPUs; otherwise, CP is the best choice. In addition, both the Intel Core i7 and the NVIDIA GTX 580 fully support the IEEE 754-2008 floating-point precision standard; hence, correctness of our GPU implementation can be fully guaranteed.