Accelerating Recurrent Neural Networks for Gravitational Wave Experiments

Accelerating Recurrent Neural Networks for Gravitational Wave Experiments
复制标题

DOI:
10.1109/asap52443.2021.00025
复制
发表时间:
2021-06
期刊:
2021 IEEE 32nd International Conference on Application-specific Systems, Architectures and Processors (ASAP)
影响因子:
--
通讯作者:
Zhiqiang Que;Erwei Wang;Umar Marikar;Eric A. Moreno;J. Ngadiuba;Hamza Javed;Bartłomiej Borzyszkowski;T. Aarrestad;V. Loncar;S. Summers;M. Pierini;P. Cheung;W. Luk
Zhiqiang Que;Erwei Wang;Umar Marikar;Eric A. Moreno;J. Ngadiuba;Hamza Javed;Bartłomiej Borzyszkowski;T. Aarrestad;V. Loncar;S. Summers;M. Pierini;P. Cheung;W. Luk
中科院分区:
其他
文献类型:
--
作者:
Zhiqiang Que;Erwei Wang;Umar Marikar;Eric A. Moreno;J. Ngadiuba;Hamza Javed;Bartłomiej Borzyszkowski;T. Aarrestad;V. Loncar;S. Summers;M. Pierini;P. Cheung;W. Luk

文献摘要

被引文献

相似文献

本文提出了一种新的可重构架构,用于减少用于探测引力波的递归神经网络(RNN)的延迟。引力干涉仪,如LIGO探测器,捕捉宇宙事件,如黑洞合并,发生在未知的时间和不同的持续时间,产生时间序列数据。我们开发了一种新的架构,能够加速RNN推理,用于分析LIGO探测器的时间序列数据。该架构基于优化多层LSTM(长短期记忆)网络中的启动间隔(II),通过为每层确定适当的重用因子。设计了一个可定制的模板,用于这种架构,它可以生成低延迟的FPGA设计,使用高层次的综合工具,有效地利用资源。基于两个LSTM模型对所提出的方法进行了评估,目标是ZYNQ 7045 FPGA和U250 FPGA。实验结果表明,平衡的II,DSP的数量可以减少到42%,而实现相同的II。与其他基于FPGA的LSTM设计相比,我们的设计可以实现约4.92至12.4倍的延迟。
This paper presents novel reconfigurable architectures for reducing the latency of recurrent neural networks (RNNs) that are used for detecting gravitational waves. Gravitational interferometers such as the LIGO detectors capture cosmic events such as black hole mergers which happen at unknown times and of varying durations, producing time-series data. We have developed a new architecture capable of accelerating RNN inference for analyzing time-series data from LIGO detectors. This architecture is based on optimizing the initiation intervals (II) in a multi-layer LSTM (Long Short-Term Memory) network, by identifying appropriate reuse factors for each layer. A customizable template for this architecture has been designed, which enables the generation of low-latency FPGA designs with efficient resource utilization using high-level synthesis tools. The proposed approach has been evaluated based on two LSTM models, targeting a ZYNQ 7045 FPGA and a U250 FPGA. Experimental results show that with balanced II, the number of DSPs can be reduced up to 42% while achieving the same IIs. When compared to other FPGA-based LSTM designs, our design can achieve about 4.92 to 12.4 times lower latency.