Accelerating Kirchhoff Migration by CPU and GPU Cooperation

Accelerating Kirchhoff Migration by CPU and GPU Cooperation
复制标题

DOI:
10.1109/sbac-pad.2009.29
复制
发表时间:
2009-10
期刊:
2009 21st International Symposium on Computer Architecture and High Performance Computing
影响因子:
--
通讯作者:
J. Panetta;T. Teixeira;P. S. Filho;C. C. Filho-C.;David Sotelo;Fernando M. Roxo da Motta;Silvio Sinedino Pinheiro-S
J. Panetta;T. Teixeira;P. S. Filho;C. C. Filho-C.;David Sotelo;Fernando M. Roxo da Motta;Silvio Sinedino Pinheiro-S
中科院分区:
其他
文献类型:
--
作者:
J. Panetta;T. Teixeira;P. S. Filho;C. C. Filho-C.;David Sotelo;Fernando M. Roxo da Motta;Silvio Sinedino Pinheiro-S

文献摘要

被引文献

相似文献

我们讨论了巴西石油公司生产克希霍夫叠前地震偏移的性能上的集群的64个GPU和256个CPU内核。将应用程序热点(单个CPU核心执行时间的98.2%)移植和优化到单个GPU,可将控制运行的总执行时间减少36倍。然后,我们反对将下一个热点(单个CPU核心执行时间的1.5%)移植到GPU的通常做法。相反,我们表明,CPU和GPU的合作减少了总执行时间的59倍,在同一个控制运行。剩余的GPU空闲周期通过使用源自不同CPU核心的多个请求使GPU过载来消除。然而,增加计算中的CPU核心数量会降低增益,这是由于在没有GPU的运行中增强的并行性和在有GPU的运行中的GPU饱和的组合。我们继续通过复制控制运行数据获得的均匀负载在整个集群上获得接近完美的加速。为了科普异构负载的真实的世界的数据,我们展示了一个动态的负载平衡方案,减少总执行时间的一个因素,20上运行,使用所有的GPU和一半的集群CPU核心的运行,使用所有的CPU核心,但没有GPU。
We discuss the performance of Petrobras production Kirchhoff prestack seismic migration on a cluster of 64 GPUs and 256 CPU cores. Porting and optimization of the application hot spot (98.2% of a single CPU core execution time) to a single GPU reduces total execution time by a factor of 36 on a control run. We then argue against the usual practice of porting the next hot spot (1.5% of single CPU core execution time) to the GPU. Instead, we show that cooperation of CPU and GPU reduces total execution time by a factor of 59 on the same control run. Remaining GPU idle cycles are eliminated by overloading the GPU with multiple requests originated from distinct CPU cores. However, increasing the number of CPU cores in the computation reduces the gain due to the combination of enhanced parallelism in the runs without GPUs and GPU saturation on runs with GPUs. We proceed by obtaining close to perfect speed-up on the full cluster over homogeneous load obtained by replicating control run data. To cope with the heterogeneous load of real world data we show a dynamic load balancing scheme that reduces total execution time by a factor of 20 on runs that use all GPUs and half of the cluster CPU cores with respect to runs that use all CPU cores but no GPU.