Parallel Sampling-Pipeline for Indefinite Stream of Heterogeneous Graphs using OpenCL for FPGAs

Parallel Sampling-Pipeline for Indefinite Stream of Heterogeneous Graphs using OpenCL for FPGAs
复制标题

DOI:
10.1109/bigdata.2018.8621979
复制
发表时间:
2018-12
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
M. Tariq;F. Saeed
M. Tariq;F. Saeed
中科院分区:
其他
文献类型:
--
作者:
M. Tariq;F. Saeed

文献摘要

相似文献

在数据科学领域,需要处理和分析大量的数据,通常表示为图形。快速有效地处理这些数据以节省时间和能源至关重要。数据的量和速度,沿着伴随着图形数据结构中的不规则访问模式,在分析和处理方面提出了挑战。此外,大量的时间和精力都花在分析大型计算集群和/或数据中心上的这些图表上。使用图形采样技术过滤和细化数据是加快分析速度的最有效方法之一。高效的加速器(如FPGA)已被证明可以显著降低运行算法的能源成本。为此,我们提出了一种并行图形采样技术的设计和实现,大量的输入图形流到FPGA中。采用了一种使用OpenCL for FPGA的并行方法,以提出一种既省时又节能的解决方案。我们介绍了一种新的图形数据结构,适合流图形的FPGA,它允许时间和内存效率的图形表示。我们的实验表明,我们提出的技术是3倍的速度和2倍的能源效率相比,串行CPU版本的算法。
In the field of data science, a huge amount of data, generally represented as graphs, needs to be processed and analyzed. It is of utmost importance that this data be processed swiftly and efficiently to save time and energy. The volume and velocity of data, along with irregular access patterns in graph data structures, pose challenges in terms of analysis and processing. Further, a big chunk of time and energy is spent on analyzing these graphs on large compute clusters and/or data-centers. Filtering and refining of data using graph sampling techniques are one of the most effective ways to speed up the analysis. Efficient accelerators, such as FPGAs, have proven to significantly lower the energy cost of running an algorithm. To this end, we present the design and implementation of a parallel graph sampling technique, for a large number of input graphs streaming into a FPGA. A parallel approach using OpenCL for FPGAs was adopted to come up with a solution that is both time- and energy-efficient. We introduce a novel graph data structure, suitable for streaming graphs on FPGAs, that allows time- and memory-efficient representation of graphs. Our experiments show that our proposed technique is 3x faster and 2x more energy efficient as compared to serial CPU version of the algorithm.