HPC Programming on Intel Many-Integrated-Core Hardware with MAGMA Port to Xeon Phi

HPC Programming on Intel Many-Integrated-Core Hardware with MAGMA Port to Xeon Phi
复制标题

在具有 MAGMA 端口至 Xeon Phi 的英特尔众集成核心硬件上进行 HPC 编程

DOI:
--
复制
发表时间:
2015
影响因子:
--
通讯作者:
S. Tomov
S. Tomov
中科院分区:
计算机科学4区
文献类型:
--
作者:
J. Dongarra;Mark Gates;A. Haidar;Yulu Jia;K. Kabir;P. Luszczek;S. Tomov

文献摘要

被引文献

相似文献

本文介绍了几种基于英特尔至强融核协处理器的多核稠密线性代数(DLA)基本算法的设计和实现。特别是,我们考虑求解线性系统的算法。此外,我们还概述了MAGMA MIC库,这是一个开源的高性能库,它结合了这里介绍的开发,并且更广泛地提供了与流行的LAPACK库等效的DLA功能,同时针对具有多核CPU和协处理器混合的异构体系结构。LAPACK兼容性简化了MAGMA MIC库在应用中的使用,同时为它们提供可移植的高性能DLA。高性能是通过使用高性能的BLAS,硬件特定的调整,和杂交方法,使我们分裂成各种粒度的计算任务的算法。通过最小化数据移动并将算法要求映射到各种异构硬件组件的架构强度,在异构硬件上适当地调度这些任务的执行。我们的方法和编程技术被整合到MAGMA MIC API中,它将应用程序开发人员从Xeon Phi架构的细节中抽象出来,因此适用于DLA范围之外的算法。
This paper presents the design and implementation of several fundamental dense linear algebra (DLA) algorithms for multicore with Intel Xeon Phi coprocessors. In particular, we consider algorithms for solving linear systems. Further, we give an overview of the MAGMA MIC library, an open source, high performance library, that incorporates the developments presented here and, more broadly, provides the DLA functionality equivalent to that of the popular LAPACK library while targeting heterogeneous architectures that feature a mix of multicore CPUs and coprocessors.The LAPACK-compliance simplifies the use of the MAGMA MIC library in applications, while providing them with portably performant DLA. High performance is obtained through the use of the high-performance BLAS, hardware-specific tuning, and a hybridization methodology whereby we split the algorithm into computational tasks of various granularities. Execution of those tasks is properly scheduled over the heterogeneous hardware by minimizing data movements and mapping algorithmic requirements to the architectural strengths of the various heterogeneous hardware components. Our methodology and programming techniques are incorporated into the MAGMA MIC API, which abstracts the application developer fromthe specifics of the Xeon Phi architecture and is therefore applicable to algorithms beyond the scope of DLA.