Cache Optimization and Performance Modeling of Batched, Small, and Rectangular Matrix Multiplication on Intel, AMD, and Fujitsu Processors

Cache Optimization and Performance Modeling of Batched, Small, and Rectangular Matrix Multiplication on Intel, AMD, and Fujitsu Processors
复制标题

Intel、AMD 和 Fujitsu 处理器上的批量、小型和矩形矩阵乘法的缓存优化和性能建模

DOI:
--
复制
发表时间:
2023
影响因子:
2.7
通讯作者:
George Bosilca
George Bosilca
中科院分区:
计算机科学3区
文献类型:
--
作者:
Sameer Deshmukh;Rio Yokota;George Bosilca

文献摘要

参考文献

相似文献

DOI: --
发表时间: 2007
影响因子: 4.3
作者:
M. Bebendorf;W. Hackbusch
通讯作者: W. Hackbusch
优化和性能模型对 CFD 中基于模板的复杂循环内核的实际适用性
DOI: 10.1177/1094342018774126
发表时间: 2019
期刊: The International Journal of High Performance Computing Applications
影响因子: --
作者:
K. Wichmann;M. Kronbichler;R. Löhner;W. Wall
通讯作者: W. Wall
GPU 上的批量平铺低阶 GEMM
DOI: --
发表时间: 2018
期刊: --
影响因子: --
作者:
A. Charara;D. Keyes;H. Ltaief
通讯作者: H. Ltaief
BLASFEO 的 BLAS API
DOI: --
发表时间: 2019
影响因子: 2.7
作者:
G. Frison;Tommaso Sartor;Andrea Zanelli;M. Diehl
通讯作者: M. Diehl
密集对称分层半可分离矩阵的分布式 O(N) 线性求解器
DOI: 10.1109/mcsoc.2019.00008
发表时间: 2019
期刊: 2019 IEEE 13th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC
影响因子: --
作者:
Yu, Chenhan D.;Reiz, Severin;Biros, George
通讯作者: Biros, George