High-Performance Tensor Contraction without Transposition

High-Performance Tensor Contraction without Transposition
复制标题

DOI:
10.1137/16m108968x
复制
发表时间:
2016-07
期刊:
SIAM J. Sci. Comput.
影响因子:
--
通讯作者:
D. Matthews
D. Matthews
中科院分区:
其他
文献类型:
--
作者:
D. Matthews

文献摘要

被引文献

相似文献

张量计算 - 尤其是张量收缩(TC)---在许多科学计算应用中是重要的内核。由于TC与矩阵乘法的基本相似性以及优化实现(例如Blas)的可用性,传统上已经根据BLAS操作实施了张量操作,从而产生了性能和存储空间。取而代之的是,我们使用柔性BLAS样的实例软件(BLIS)框架实现TC,该框架允许将张量的换位(重塑)与内部分区和包装操作融合,不需要明确的转置操作或其他工作空间。这种实现,TBLI,实现了矩阵乘法的性能,在某些情况下,其绩效高于传统TC的性能。我们的实现支持多线程使用与BLIS中矩阵乘法相同的方法,具有相似的性能特征。复杂性...
Tensor computations---in particular tensor contraction (TC)---are important kernels in many scientific computing applications. Due to the fundamental similarity of TC to matrix multiplication and to the availability of optimized implementations such as the BLAS, tensor operations have traditionally been implemented in terms of BLAS operations, incurring both a performance and a storage overhead. Instead, we implement TC using the flexible BLAS-like Instantiation Software (BLIS) framework, which allows for transposition (reshaping) of the tensor to be fused with internal partitioning and packing operations, requiring no explicit transposition operations or additional workspace. This implementation, TBLIS, achieves performance approaching that of matrix multiplication, and in some cases considerably higher than that of traditional TC. Our implementation supports multithreading using an approach identical to that used for matrix multiplication in BLIS, with similar performance characteristics. The complexity...