High Performance Implementation of RTM Seismic Modeling on FPGAs: Architecture, Arithmetic and Power Issues

High Performance Implementation of RTM Seismic Modeling on FPGAs: Architecture, Arithmetic and Power Issues
复制标题

DOI:
10.1007/978-1-4614-1791-0_10
复制
发表时间:
2013
影响因子:
1
通讯作者:
V. Medeiros;A. C. Barros;A. Silva-Filho;M. Lima
V. Medeiros;A. C. Barros;A. Silva-Filho;M. Lima
中科院分区:
--
文献类型:
--
作者:
V. Medeiros;A. C. Barros;A. Silva-Filho;M. Lima

文献摘要

被引文献

相似文献

本文介绍了石油天然气行业的一个实例,即二维逆时偏移(RTM)地震建模算法的FPGA实现。这些设备主要用作科学计算应用中的加速器,这些应用需要大量数据处理,大型并行机,巨大的内存带宽和功率。RTM算法使您能够直接精确地解决复杂地质结构中的声波和弹性波问题,需要高计算能力。面对这样的挑战,我们建议的策略,如降低算术精度,基于定点数,并建议一个高度并行的架构。在本章中,通过信噪比(SRN)和通用图像质量指数(UIQI)指标分析了存储/处理数据的精度降低的影响。结果表明,对于15位字长的偏移图像,SRN大于50 dB可以被认为是可以接受的。一个特殊的流处理架构,旨在实现最佳的数据重用的算法。它是通过FPGA内部存储器中基于FIFO的高速缓存来实现的。还开发了时间流水线结构,允许同时执行多个时间步骤。这种方法的主要优点是能够保持处理一个时间步长所需的相同内存带宽。同时处理的时间步数受到FPGA内部存储器和逻辑块数量的限制。该算法在Altera Stratix 260 E上实现,具有16个处理元件(PE)。FPGA比CPU快29倍,比GPGPU慢13%。在功耗方面,CPU+FPGA的效率是GPGPU系统的1.7倍。
This work presents a case study in the oil and gas industry, namely the FPGA implementation of the 2D reverse timing migration (RTM) seismic modeling algorithm. These devices have been largely used as accelerators in scientific computing applications that require massive data processing, large parallel machines, huge memory bandwidth and power. The RTM algorithm enables you to directly solve the acoustic and elastic waves problems with precision in complex geological structures, demanding a high computational power. To face such challenges we suggest strategies such as reduced arithmetic precision, based on fixed-point numbers, and a highly parallel architecture are suggested. The effects of such reduced precision for storage/processing data are analyzed in this chapter through signal-noise ratio (SRN) and universal image quality index (UIQI) metrics. The results show that SRN higher than 50dB can be considered acceptable for a migrated image with 15 bits word size. A special stream-processing architecture aiming to implement the best possible data reuse for the algorithm is also presented. It was implemented by an FIFO-based cache in the internal memory of the FPGA. A temporal pipeline structure has also been developed, allowing that multiple time steps to be performed at the same time. The main advantage of this approach is the ability to keep the same memory bandwidth needs of processing just one time step. The number of time steps processed at the same time is limited by the amount of FPGA internal memory and logic blocks. The algorithm was implemented on an Altera Stratix 260E, with 16 processing elements (PEs). The FPGA was 29 times faster than the CPU and only 13% slower than the GPGPU. In terms of power consumption, the CPU+FPGA was 1.7 times more efficient than the GPGPU system.