Memory-efficient parallel tensor decompositions

Memory-efficient parallel tensor decompositions
复制标题

内存高效的并行张量分解

DOI:
10.1109/hpec.2017.8091026
复制
发表时间:
2017
期刊:
2017 IEEE High Performance Extreme Computing Conference (HPEC)
影响因子:
--
通讯作者:
R. Lethin
R. Lethin
中科院分区:
--
文献类型:
--
作者:
M. Baskaran;Thomas Henretty;B. Pradelle;M. H. Langston;David Bruns;J. Ezick;R. Lethin

文献摘要

被引文献

相似文献

张量分解是一种强大的技术,可以对真实世界的数据进行全面和完整的分析。通过张量分解进行数据分析涉及到大规模不规则稀疏数据的密集计算。优化此类数据密集型计算的执行是减少现实世界数据分析应用程序中的解决方案时间(或响应时间)的关键。随着高性能计算(HPC)系统越来越多地被用于数据分析应用,优化稀疏张量计算并在现代先进的HPC系统上高效执行变得越来越重要。除了利用HPC系统的大处理能力外,提高HPC系统中的内存性能(内存使用率、通信、同步、内存重用和数据局部性)也是至关重要的。在这篇文章中,我们提出了多个优化,目标是在高性能计算系统上更快、更高效地执行大规模张量分析。我们证明,当我们的技术应用于来自不同应用领域的不同大小和结构的多个数据集时,我们的技术可以减少张量分解方法的内存使用量和执行时间。我们的内存使用量最多减少到原来的1/11,性能提高多达1/7。更重要的是,我们能够在多核系统上对一些重要的数据集应用大型张量分解,如果没有我们的优化,这是不可行的。
Tensor decompositions are a powerful technique for enabling comprehensive and complete analysis of real-world data. Data analysis through tensor decompositions involves intensive computations over large-scale irregular sparse data. Optimizing the execution of such data intensive computations is key to reducing the time-to-solution (or response time) in real-world data analysis applications. As high-performance computing (HPC) systems are increasingly used for data analysis applications, it is becoming increasingly important to optimize sparse tensor computations and execute them efficiently on modern and advanced HPC systems. In addition to utilizing the large processing capability of HPC systems, it is crucial to improve memory performance (memory usage, communication, synchronization, memory reuse, and data locality) in HPC systems. In this paper, we present multiple optimizations that are targeted towards faster and memory-efficient execution of large-scale tensor analysis on HPC systems. We demonstrate that our techniques achieve reduction in memory usage and execution time of tensor decomposition methods when they are applied on multiple datasets of varied size and structure from different application domains. We achieve up to 11× reduction in memory usage and up to 7× improvement in performance. More importantly, we enable the application of large tensor decompositions on some important datasets on a multi-core system that would not have been feasible without our optimization.