Exploiting Data Sparsity for Large-Scale Matrix Computations

Exploiting Data Sparsity for Large-Scale Matrix Computations
复制标题

利用数据稀疏性进行大规模矩阵计算

DOI:
--
复制
发表时间:
2018
期刊:
European Conference on Parallel Processing
影响因子:
--
通讯作者:
D. Keyes
D. Keyes
中科院分区:
--
文献类型:
--
作者:
Kadir Akbudak;H. Ltaief;A. Mikhalev;A. Charara;Aniello Esposito;D. Keyes

文献摘要

被引文献

相似文献

利用密集矩阵中的数据稀疏性是架构之间的算法桥梁,这些架构在每个核心的基础上内存越来越简朴,与极端规模的应用程序之间。在这项工作中,我们利用多核架构上的分层矩阵计算 (HiCMA) 库,通过显着减少求解时间和内存占用来解决这一具有挑战性的问题,同时保留应用程序的指定精度要求。我们扩展了 HiCMA,以在分布式内存系统上提供大规模科学应用中最广泛使用的矩阵分解之一(即 Cholesky 分解)的高性能实现。它采用平铺低秩数据格式来压缩矩阵的密集数据稀疏非对角平铺。然后,它将矩阵计算分解为相互依赖的任务,并依靠动态运行时系统 StarPU 进行异步乱序调度,同时允许高用户生产力。高达 1100 万矩阵维度上的性能比较和内存占用表明,与最先进的开源和供应商优化的数值库相比,数千个内核上的这两个指标的性能增益和内存节省都超过一个数量级。这是实现大规模矩阵计算以解决气候/天气预报应用的地理空间统计中的大数据问题的一个重要里程碑。
Exploiting data sparsity in dense matrices is an algorithmic bridge between architectures that are increasingly memory-austere on a per-core basis and extreme-scale applications. In this work, we leverage the Hierarchical matrix Computations on Manycore Architectures (HiCMA) library in order to tackle this challenging problem by achieving significant reductions in time to solution and memory footprint, while preserving a specified accuracy requirement of the application. We have extended HiCMA to provide a high-performance implementation on distributed-memory systems of one of the most widely used matrix factorization in large-scale scientific applications, i.e., the Cholesky factorization. It employs the tile low-rank data format to compress the dense data-sparse off-diagonal tiles of the matrix. It then decomposes the matrix computations into interdependent tasks and relies on the dynamic runtime system StarPU for asynchronous out-of-order scheduling, while allowing high user productivity. Performance comparisons and memory footprint on matrix dimensions up to eleven million show a performance gain and memory saving of more than an order of magnitude for both metrics on thousands of cores, against state-of-the-art open-source and vendor optimized numerical libraries. This represents an important milestone in enabling large-scale matrix computations toward solving big data problems in geospatial statistics for climate/weather forecasting applications.