Performance Evaluation of Low Level Multithreaded BLAS Kernels on Intel Processor Based cc-NUMA Systems

Performance Evaluation of Low Level Multithreaded BLAS Kernels on Intel Processor Based cc-NUMA Systems
复制标题

DOI:
10.1007/978-3-540-39707-6_45
复制
发表时间:
2003-10
期刊:
--
影响因子:
--
通讯作者:
A. Nishida;Y. Oyanagi
A. Nishida;Y. Oyanagi
中科院分区:
其他
文献类型:
--
作者:
A. Nishida;Y. Oyanagi

文献摘要

相似文献

计算线性代数中稀疏矩阵算法的BLAS库的并行实现是一个关键问题,特别是在有限存储带宽的共享存储器结构上。在这项研究中,我们使用低级别的多线程BLAS内核的CC-NUMA系统的性能进行评估。在两个基于Intel处理器的体系结构NEC TX 7/AzusA和IBM xSeries 440上对编译器和系统的性能进行了评估。
Parallel implementation of the BLAS library for sparse matrix algorithms in computational linear algebra is a critical problem, especially on the shared memory architectures with finite memory bandwidth. In this study, we evaluate the performance of the cc-NUMA systems using low level multithreaded BLAS kernels. The performance of both the compiler and the systems are evaluated on two Intel processor based architectures, NEC TX7/AzusA and IBM xSeries 440.