Kokkos Array performance-portable manycore programming model

Kokkos Array performance-portable manycore programming model
复制标题

Kokkos Array 高性能便携式众核编程模型

DOI:
--
复制
发表时间:
2012
期刊:
Programming Models and Applications for Multicores and Manycores
影响因子:
--
通讯作者:
Daniel Sunderland
Daniel Sunderland
中科院分区:
--
文献类型:
--
作者:
H. C. Edwards;Daniel Sunderland

文献摘要

被引文献

相似文献

大型、复杂的科学和工程应用代码在实现其数学模型的计算内核上投入了大量资金。考虑到不同的编程模型、应用程序编程接口 (API) 和性能要求,将这些计算内核移植到多核 CPU 和众核加速器(例如 NVIDIA® GPU)设备是一项重大挑战。 Kokkos 阵列编程模型提供了基于库的方法来实现计算内核,这些内核的性能可移植到多核 CPU 和众核加速器设备。该编程模型基于三个基本概念:(1) 每个都有自己的内存空间的众核计算设备,(2) 数据并行计算内核,以及 (3) 多维数组。通过直观的多维数组 API 将计算内核与设备特定的数据访问性能要求(例如 NVIDIA 合并内存访问)解耦,从而实现性能可移植性。 Kokkos Array API 使用 C++ 模板元编程在编译时将设备最佳数据访问映射透明地插入到计算内核中。利用这种编程模型,计算内核可以编写一次,无需修改,即可性能可移植地编译为多核 CPU 和众核加速器设备。
Large, complex scientific and engineering application code have a significant investment in computational kernels which implement their mathematical models. Porting these computational kernels to multicore-CPU and manycore-accelerator (e.g., NVIDIA® GPU) devices is a major challenge given the diverse programming models, application programming interfaces (APIs), and performance requirements. The Kokkos Array programming model provides library-based approach for implementing computational kernels that are performance-portable to multicore-CPU and manycore-accelerator devices. This programming model is based upon three fundamental concepts: (1) manycore compute devices each with its own memory space, (2) data parallel computational kernels, and (3) multidimensional arrays. Performance-portability is achieved by decoupling computational kernels from device-specific data access performance requirements (e.g., NVIDIA coalesced memory access) through an intuitive multidimensional array API. The Kokkos Array API uses C++ template meta-programming to, at compile time, transparently insert device-optimal data access maps into computational kernels. With this programming model computational kernels can be written once and, without modification, performance-portably compiled to multicore-CPU and manycore-accelerator devices.