Developing an Efficient Vector-Friendly Implementation of the Breadth-First Search Algorithm for NEC SX-Aurora TSUBASA

Developing an Efficient Vector-Friendly Implementation of the Breadth-First Search Algorithm for NEC SX-Aurora TSUBASA
复制标题

为 NEC SX-Aurora TSUBASA 开发广度优先搜索算法的高效向量友好实现

DOI:
10.1007/978-3-030-55326-5_10
复制
发表时间:
2020
期刊:
Parallel Computational Technologies. PCT 2020. Communications in Computer and Information Science
影响因子:
--
通讯作者:
Kobayashi Hiroaki
Kobayashi Hiroaki
中科院分区:
--
文献类型:
--
作者:
Afanasyev Ilya V.;Voevodin Vladimir V.;Komatsu Kazuhiko;Kobayashi Hiroaki

文献摘要

参考文献

相似文献

广度优先搜索是一种重要的计算核心,是许多图算法的基础。到目前为止,已经针对各种计算平台提出了不同的算法和实现方法来解决BFS问题,其中方向优化算法对于许多现实世界的图形类型来说是最快和最有效的。然而,由于图数据结构和算法本身的高度不规则性,直接实现向量计算机的方向优化BFS可能极具挑战性和效率低下。本文介绍了世界上第一次尝试,旨在创建一个有效的矢量友好的BFS实现的方向优化算法的NEC SX极光TSUBASA架构。SX-Aurora TSUBASA矢量处理器提供高性能计算能力以及世界上最高的带宽内存,使其成为解决各种图形处理问题的非常有趣的平台。本文中提出的实现方式明显优于现代CPU(Intel Skylake)和NVIDIA V100 GPU的现有最先进的实现方式。此外,与其他平台和实现方式相比,所提出的实现方式在平均功耗和实现的每瓦性能方面都实现了显著更高的能效。
Breadth-First Search (BFS) is an important computational kernel used as a building-block for many other graph algorithms. Different algorithms and implementation approaches aimed to solve the BFS problem have been proposed so far for various computational platforms, with the direction-optimizing algorithm being the fastest and the most computationally efficient for many real-world graph types. However, straightforward implementation of direction-optimizing BFS for vector computers can be extremely challenging and inefficient due to the high irregularity of graph data structure and the algorithm itself. This paper describes the world’s first attempt aimed to create an efficient vector-friendly BFS implementation of the direction-optimizing algorithm for NEC SX-Aurora TSUBASA architecture. SX-Aurora TSUBASA vector processors provide high-performance computational power together with a world-highest bandwidth memory, making it a very interesting platform for solving various graph-processing problems. The implementation proposed in this paper significantly outperforms the existing state-of-the-art implementations both for modern CPUs (Intel Skylake) and NVIDIA V100 GPUs. In addition, the proposed implementation achieves significantly higher energy efficiency compared to other platforms and implementations both in terms of average power consumption and achieved performance per watt.
GPU 上的高效混合广度优先搜索
DOI: 10.1007/978-3-319-03889-6_5
发表时间: 2013
期刊: --
影响因子: --
作者:
Takaaki Hiragushi;D. Takahashi
通讯作者: D. Takahashi
NVIDIA GPU 与 NEC SX-Aurora TSUBASA 矢量处理器中使用的 SIMD 处理功能之间的关系分析
DOI: 10.1007/978-3-030-25636-4_10
发表时间: 2019
影响因子: 5.3
作者:
I. Afanasyev;V. Voevodin;V. Voevodin;Kazuhiko Komatsu;Hiroaki Kobayashi
通讯作者: Hiroaki Kobayashi
为 NEC SX-ACE 开发 Bellman-Ford 和前向后向图算法的高效实现
DOI: 10.14529/jsfi180311
发表时间: 2018
期刊: Supercomput. Front. Innov.
影响因子: --
作者:
I. Afanasyev;A. Antonov;D. Nikitenko;Vadim V. Voevodin;V. Voevodin;Kazuhiko Komatsu;Osamu Watanabe;A. Musa;Hiroaki Kobayashi
通讯作者: Hiroaki Kobayashi