Construction of Performance Model of Tile CAQR and Performance Result of the Implementation

Construction of Performance Model of Tile CAQR and Performance Result of the Implementation
复制标题

Tile CAQR性能模型构建及实施性能结果

DOI:
10.1109/mcsoc.2017.18
复制
发表时间:
2017
期刊:
2017 IEEE 11th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)
影响因子:
--
通讯作者:
Tomohiro Suzuki
Tomohiro Suzuki
中科院分区:
--
文献类型:
--
作者:
M. Takayanagi;Tomohiro Suzuki

文献摘要

被引文献

相似文献

可以通过异步执行许多细粒度任务来利用高度并行的计算资源。矩阵分解的tile算法可以生成许多细粒度的任务,因此适合现代多核/多核架构。然而,该算法的性能在很大程度上取决于贴图的大小。我们在集群系统上以OpenMP/MPI混合方式实现了平铺算法,并构建了一个性能模型,该模型通过测量实现中简单计算内核的性能来调整平铺大小。在本报告中,我们在K计算机上测试了我们的高矩阵和瘦矩阵的避免通信的平铺QR实现,并演示了性能模型的适用性。
Highly parallel computational resources can be exploited by asynchronously executing many fine-grained tasks. The tile algorithm for matrix decomposition can generate many fine-grained tasks, so is suitable for modern multicore/manycore architectures. However, the performance of this algorithm significantly depends on the tile size. We implement the tile algorithm in OpenMP/MPI hybrid fashion on a cluster system and construct a performance model that tunes the tile size by measuring the performance of simple computational kernels in our implementation. In this report, we test our communication-avoiding tile QR implementation for tall and skinny matrices on the K computer, and demonstrate the applicability of the performance model.