neoSYCL: a SYCL implementation for SX-Aurora TSUBASA

neoSYCL: a SYCL implementation for SX-Aurora TSUBASA
复制标题

neoSYCL:SX-Aurora TSUBASA 的 SYCL 实现

DOI:
10.1145/3432261.3432268
复制
发表时间:
2021
期刊:
International Conference on High Performance Computing in Asia-Pacific Region (HPC Asia 2021)
影响因子:
--
通讯作者:
Mulya Agung and Hiroyuki Takizawa
Mulya Agung and Hiroyuki Takizawa
中科院分区:
--
文献类型:
--
作者:
Yinan Ke;Mulya Agung and Hiroyuki Takizawa

文献摘要

参考文献

被引文献

相似文献

最近,高性能计算世界已经转向更加异构的架构。因此,将部分应用程序执行卸载到专用加速器已成为标准做法。然而,生产力方面的劣势仍然是加速器编程的一个问题。本文提出了 neoSYCL:SX-Aurora TSUBASA 的 SYCL 实现,旨在提高生产力并实现与本机实现相当的性能。与其他实现不同,neoSYCL 可以在源代码级别识别和分离 SYCL 代码的内核部分。因此,可以使用卸载编程模型轻松地将这种方法转移到任何异构架构。在本文中,我们展示了对 SX-Aurora TSUBASA 的评估结果。为了定量讨论性能和生产力,我们使用两种不同的基准和代码复杂性指标进行评估。结果表明,neoSYCL 可以提高生产力,同时达到与本机实现相同的性能。
Recently, the high-performance computing world has moved to more heterogeneous architectures. Thus, it has become a standard practice to offload a part of application execution to dedicated accelerators. However, the disadvantage in productivity is still a problem in programming for accelerators. This paper proposes neoSYCL: a SYCL implementation for SX-Aurora TSUBASA, aiming to improve productivity and achieve comparable performance with native implementations. Unlike other implementations, neoSYCL can identify and separate the kernel part of the SYCL code at the source code level. Thus, this approach can easily be moved to any heterogeneous architectures using the offload programming model. In this paper, we show the evaluation results on SX-Aurora TSUBASA. To quantitatively discuss not only performance but also the productivity, we use two different benchmarks and code-complexity metrics for the evaluation. The results show that neoSYCL can improve productivity while reaching the same performance as native implementations.
NEC SX-Aurora TSUBASA 上用于卸载的异构活动消息
DOI: 10.1109/ipdpsw.2019.00014
发表时间: 2019
期刊: 2019 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
影响因子: --
作者:
M. Noack;E. Focht;T. Steinke
通讯作者: T. Steinke
C 语言的反射系统
DOI: --
发表时间: 2004
期刊: ArXiv
影响因子: --
作者:
Duraid Madina;R. Standish
通讯作者: R. Standish
SX-Aurora TSUBASA 的 I/O 性能
DOI: 10.1109/ipdpsw50202.2020.00014
发表时间: 2020
期刊: Proceedings of 2020 IEEE International Parallel and Distributed Processing Symposium Workshop (IPDPSW)
影响因子: --
作者:
Yokokawa Mitsuo;Nakai Ayano;Komatsu Kazuhiko;Watanabe Yuta;Masaoka Yasuhisa;Isobe Yoko;Kobayashi Hiroaki
通讯作者: Kobayashi Hiroaki
GPU-STREAM v2.0:跨多种并行编程模型对多核处理器可实现的内存带宽进行基准测试
DOI: --
发表时间: 2016
期刊: ISC Workshops
影响因子: --
作者:
Tom Deakin;J. Price;Matt Martineau;Simon McIntosh
通讯作者: Simon McIntosh