Implementing Multifrontal Sparse Solvers for Multicore Architectures with Sequential Task Flow Runtime Systems

Implementing Multifrontal Sparse Solvers for Multicore Architectures with Sequential Task Flow Runtime Systems
复制标题

使用顺序任务流运行时系统实现多核架构的多前沿稀疏求解器

DOI:
10.1145/2898348
复制
发表时间:
2016
期刊:
ACM Transactions on Mathematical Software (TOMS)
影响因子:
--
通讯作者:
Florent Lopez
Florent Lopez
中科院分区:
--
文献类型:
--
作者:
E. Agullo;A. Buttari;A. Guermouche;Florent Lopez

文献摘要

被引文献

相似文献

面对多核处理器的出现和硬件架构日益复杂的情况,基于DAG并行的编程模型在高性能科学计算社区重新流行起来。现代运行时系统提供了一个符合这种范式的编程接口和强大的引擎,用于调度应用程序分解成的任务。这些工具已经证明了它们在一些稠密线性代数应用中的有效性。本文评估了基于复杂应用程序的顺序任务流模型的运行时系统的可用性和有效性,即稀疏矩阵多前沿分解,具有极不规则的工作负载,具有不同粒度和特征的任务,以及可变的内存消耗。最重要的是,它展示了这种并行编程模型如何简化复杂功能的开发,这些功能有利于稀疏直接求解器的性能及其内存消耗。我们用运行在StarPU运行时系统之上的多前端QR分解来说明我们的讨论。
To face the advent of multicore processors and the ever increasing complexity of hardware architectures, programming models based on DAG parallelism regained popularity in the high performance, scientific computing community. Modern runtime systems offer a programming interface that complies with this paradigm and powerful engines for scheduling the tasks into which the application is decomposed. These tools have already proved their effectiveness on a number of dense linear algebra applications. This article evaluates the usability and effectiveness of runtime systems based on the Sequential Task Flow model for complex applications, namely, sparse matrix multifrontal factorizations that feature extremely irregular workloads, with tasks of different granularities and characteristics and with a variable memory consumption. Most importantly, it shows how this parallel programming model eases the development of complex features that benefit the performance of sparse, direct solvers as well as their memory consumption. We illustrate our discussion with the multifrontal QR factorization running on top of the StarPU runtime system.