Communication avoiding and overlapping for numerical linear algebra

Communication avoiding and overlapping for numerical linear algebra
复制标题

数值线性代数的通信避免和重叠

DOI:
10.1109/sc.2012.32
复制
发表时间:
2012
期刊:
2012 International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
K. Yelick
K. Yelick
中科院分区:
--
文献类型:
--
作者:
E. Georganas;J. González;Edgar Solomonik;Yili Zheng;J. Touriño;K. Yelick

文献摘要

被引文献

相似文献

为了有效地将稠密线性代数问题扩展到未来的百亿级系统,必须避免或重叠通信成本。避免通信的2.5D算法通过以额外内存使用为代价减少处理器间数据传输量来提高可扩展性。通信重叠试图通过流水线消息和重叠计算工作来隐藏消息传递延迟。我们研究了这两种技术的相互作用和兼容性的两个矩阵乘法算法(加农和SUMMA),三角求解,和Cholesky分解。对于每个算法,我们构建了一个详细的性能模型,同时考虑关键路径依赖性和空闲时间。我们给出了新的实现2.5D算法与重叠的这些问题。我们的软件采用UPC,分区全局地址空间(PGAS)语言,提供快速的单边通信。我们表明,通信避免和重叠提供了累积的好处,核心计数规模,包括使用超过24K核心的Cray XE6系统的结果。
To efficiently scale dense linear algebra problems to future exascale systems, communication cost must be avoided or overlapped. Communication-avoiding 2.5D algorithms improve scalability by reducing inter-processor data transfer volume at the cost of extra memory usage. Communication overlap attempts to hide messaging latency by pipelining messages and overlapping with computational work. We study the interaction and compatibility of these two techniques for two matrix multiplication algorithms (Cannon and SUMMA), triangular solve, and Cholesky factorization. For each algorithm, we construct a detailed performance model that considers both critical path dependencies and idle time. We give novel implementations of 2.5D algorithms with overlap for each of these problems. Our software employs UPC, a partitioned global address space (PGAS) language that provides fast one-sided communication. We show communication avoidance and overlap provide a cumulative benefit as core counts scale, including results using over 24K cores of a Cray XE6 system.