Parallel Algorithms for Tensor Train Arithmetic

Parallel Algorithms for Tensor Train Arithmetic
复制标题

DOI:
10.1137/20m1387158
复制
发表时间:
2020-11
期刊:
SIAM J. Sci. Comput.
影响因子:
--
通讯作者:
Hussam Al Daas;Grey Ballard;P. Benner
Hussam Al Daas;Grey Ballard;P. Benner
中科院分区:
其他
文献类型:
--
作者:
Hussam Al Daas;Grey Ballard;P. Benner

文献摘要

相似文献

我们提出了高效和可扩展的并行算法,用于执行数学运算的低秩张量表示的张量列车(TT)格式。我们认为算法的加法,元素乘法,计算规范和内积,正交化,舍入(秩截断)。这些是应用程序的核心操作,例如利用TT结构的迭代Krylov求解器。并行算法是为分布式内存计算而设计的,我们使用数据分布和策略,在TT格式内为各个核心并行计算。我们分析了所提出的算法的计算和通信成本,以显示其可扩展性,我们提出了数值实验,证明其效率的共享内存和分布式内存并行系统。例如,我们在舍入2GB TT张量时观察到比现有MATLAB TT-100更好的单核性能,并且我们的实现使用单个节点的所有40个核心实现了34\times $加速。我们还展示了在所有数学运算中,在高达10,000多个核心的较大TT张量上的近线性并行缩放。
We present efficient and scalable parallel algorithms for performing mathematical operations for low-rank tensors represented in the tensor train (TT) format. We consider algorithms for addition, elementwise multiplication, computing norms and inner products, orthogonalization, and rounding (rank truncation). These are the kernel operations for applications such as iterative Krylov solvers that exploit the TT structure. The parallel algorithms are designed for distributed-memory computation, and we use a data distribution and strategy that parallelizes computations for individual cores within the TT format. We analyze the computation and communication costs of the proposed algorithms to show their scalability, and we present numerical experiments that demonstrate their efficiency on both shared-memory and distributed-memory parallel systems. For example, we observe better single-core performance than the existing MATLAB TT-Toolbox in rounding a 2GB TT tensor, and our implementation achieves a $34\times$ speedup using all 40 cores of a single node. We also show nearly linear parallel scaling on larger TT tensors up to over 10,000 cores for all mathematical operations.