High-performance implementation of the level-3 BLAS

High-performance implementation of the level-3 BLAS
复制标题

DOI:
10.1145/1377603.1377607
复制
发表时间:
2008-07-01
影响因子:
2.7
通讯作者:
Van De Geijn, Robert
Van De Geijn, Robert
中科院分区:
计算机科学3区
文献类型:
--
作者:
Goto, Kazushige;Van De Geijn, Robert

文献摘要

被引文献

相似文献

提出了一种简单但高效的方法,用于将基于高速缓存的矩阵-矩阵乘法架构上的高性能实现转换为其他常用的矩阵-矩阵计算(3级BLAS)的实现。卓越的性能在各种架构上得到了证明。
A simple but highly effective approach for transforming high-performance implementations on cache-based architectures of matrix-matrix multiplication into implementations of other commonly used matrix-matrix computations (the level-3 BLAS) is presented. Exceptional performance is demonstrated on various architectures.