Large-Scale Data Computing Performance Comparisons on SYCL Heterogeneous Parallel Processing Layer Implementations

Large-Scale Data Computing Performance Comparisons on SYCL Heterogeneous Parallel Processing Layer Implementations
复制标题

SYCL异构并行处理层实现的大规模数据计算性能比较

DOI:
10.3390/app10051656
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Nakhoon Baek
Nakhoon Baek
中科院分区:
--
文献类型:
--
作者:
Woosuk Shin;K. Yoo;Nakhoon Baek

文献摘要

被引文献

相似文献

如今,许多大数据应用程序需要大规模并行任务来计算复杂的数学运算。为了执行并行任务,像CUDA(计算统一设备架构)和OpenCL(开放计算语言)这样的平台被广泛使用和开发,以提高大规模并行任务的吞吐量。还需要对这些大规模并行计算平台进行高级抽象和平台独立性。最近,Khronos集团宣布了SYCL(C++ Single-source Heterogeneous Programming for OpenCL),一个新的跨平台抽象层,为单源异构计算提供了一种有效的方法,具有C++模板级抽象。然而,由于SYCL还没有正式的实现,我们目前有几个来自不同供应商的不同实现。在本文中,我们分析了这些SYCL实现的特点。我们还展示了这些SYCL实现的性能指标,特别是对于众所周知的大规模并行任务。我们表明,每种实现在计算不同类型的数学运算以及不同大小的数据方面都有自己的优势,沿着。我们的分析可用于大规模并行计算的抽象级成本效益使用的基本测量,特别是对于大数据应用程序。
Today, many big data applications require massively parallel tasks to compute complicated mathematical operations. To perform parallel tasks, platforms like CUDA (Compute Unified Device Architecture) and OpenCL (Open Computing Language) are widely used and developed to enhance the throughput of massively parallel tasks. There is also a need for high-level abstractions and platform-independence over those massively parallel computing platforms. Recently, Khronos group announced SYCL (C++ Single-source Heterogeneous Programming for OpenCL), a new cross-platform abstraction layer, to provide an efficient way for single-source heterogeneous computing, with C++-template-level abstractions. However, since there has been no official implementation of SYCL, we currently have several different implementations from various vendors. In this paper, we analyse the characteristics of those SYCL implementations. We also show performance measures of those SYCL implementations, especially for well-known massively parallel tasks. We show that each implementation has its own strength in computing different types of mathematical operations, along with different sizes of data. Our analysis is available for fundamental measurements of the abstract-level cost-effective use of massively parallel computations, especially for big-data applications.