Performance Evaluation of Low Level Multithreaded BLAS Kernels on Intel Processor Based cc-NUMA Systems
Performance Evaluation of Low Level Multithreaded BLAS Kernels on Intel Processor Based cc-NUMA Systems
复制标题
DOI:
10.1007/978-3-540-39707-6_45
复制
发表时间:
2003-10
期刊:
影响因子:
--
通讯作者:
A. Nishida;Y. Oyanagi
中科院分区:
文献类型:
--
作者:
A. Nishida;Y. Oyanagi
Parallel implementation of the BLAS library for sparse matrix algorithms in computational linear algebra is a critical problem, especially on the shared memory architectures with finite memory bandwidth. In this study, we evaluate the performance of the cc-NUMA systems using low level multithreaded BLAS kernels. The performance of both the compiler and the systems are evaluated on two Intel processor based architectures, NEC TX7/AzusA and IBM xSeries 440.