Analysis of Relationship Between SIMD-Processing Features Used in NVIDIA GPUs and NEC SX-Aurora TSUBASA Vector Processors

Analysis of Relationship Between SIMD-Processing Features Used in NVIDIA GPUs and NEC SX-Aurora TSUBASA Vector Processors
复制标题

NVIDIA GPU 与 NEC SX-Aurora TSUBASA 矢量处理器中使用的 SIMD 处理功能之间的关系分析

DOI:
10.1007/978-3-030-25636-4_10
复制
发表时间:
2019
影响因子:
5.3
通讯作者:
Hiroaki Kobayashi
Hiroaki Kobayashi
中科院分区:
计算机科学2区
文献类型:
--
作者:
I. Afanasyev;V. Voevodin;V. Voevodin;Kazuhiko Komatsu;Hiroaki Kobayashi

文献摘要

被引文献

相似文献

本文全面分析了三种高性能架构的主要SIMD处理特性和计算特性:两种NVIDIA GPU架构(Pascal和Volta两代)和NEC SX Aurora TSUBASA矢量处理器。由于这两种类型的架构都强烈依赖于使用SIMD处理功能,因此可以在它们之间找到数据处理原理的某些相似之处。然而,尽管在NVIDIA GPU和NEC SX-Aurora TSUBASA架构中都包含了矢量化数据处理,但这两种架构的矢量化功能是以完全不同的方式实现的。这些差异导致了几个基本的限制类的算法,可以有效地实现相应的平台。本文致力于研究移植各种类别的程序和算法之间讨论的架构,重点是利用所有可用的矢量化功能的可能性。然而,如果不详细分析这些架构中相似和不同的SIMD处理功能,就不可能解决这个问题。所进行的分析使我们能够确定几个重要的典型应用程序和算法的例子。其中一些在NVIDIA GPU和NEC SX-Aurora TSUBASA矢量处理器上表现出相当的效率,而另一些则表现出不同的效率,包括归约操作,依赖于频繁间接内存访问的程序以及通过协处理器互连的数据传输。此外,进行的分析,可以很容易地扩展这组例子来解决问题的自动移植程序之间的审查架构,我们认为我们未来的研究的一个重要方向。
This paper presents comprehensive analysis of main SIMD-processing features and computational characteristics of three high performance architectures: two NVIDIA GPU architectures (of Pascal and Volta generations) and NEC SX-Aurora TSUBASA vector processor. Since both these types of architectures strongly rely on using SIMD-processing features, certain similarities of data-processing principles can be found between them. However, despite having vectorised data-processing included in both NVIDIA GPU and NEC SX-Aurora TSUBASA architectures, vectorisation features of both architectures are implemented in completely different ways. These differences lead to several fundamental restrictions on classes of algorithms which can be efficiently implemented on corresponding platforms. This paper is devoted to the research of the possibility of porting various classes of programs and algorithms among the discussed architectures with a focus on utilising all vectorisation features available. However, without a detailed analysis of similar and different SIMD-processing features in these architectures, it is impossible to approach this problem. The performed analysis allowed us to identify several important examples of typical applications and algorithms. Some of them demonstrated comparable and the others showed different efficiency on NVIDIA GPUs and NEC SX-Aurora TSUBASA vector processors, including reduction operations, programs relying on frequent indirect memory accesses and data-transfers through co-processor interconnect. Moreover, the conducted analysis allows to easily extend this set of examples to approach the problem of automated porting of programs between the reviewed architectures, what we consider as an important direction of our future research.