Anatomy of high-performance matrix multiplication

Anatomy of high-performance matrix multiplication
复制标题

DOI:
10.1145/1356052.1356053
复制
发表时间:
2008-01-01
影响因子:
2.7
通讯作者:
Van De Geijn, Robert A.
Van De Geijn, Robert A.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Goto, Kazushige;Van De Geijn, Robert A.

文献摘要

被引文献

相似文献

我们提出了矩阵矩阵乘法的高性能实现的基本原则,这是广泛使用的GotoBLAS库的一部分。设计决策是合理的,通过不断完善的模型与多级存储器的架构。一个简单但有效的算法执行此操作的结果。广泛的选择架构上的实现实现,以实现近峰值的性能。
We present the basic principles that underlie the high-performance implementation of the matrix-matrix multiplication that is part of the widely used GotoBLAS library. Design decisions are justified by successively refining a model of architectures with multilevel memories. A simple but effective algorithm for executing this operation results. Implementations on a broad selection of architectures are shown to achieve near-peak performance.