Acceleration of LU decomposition supporting double-double, triple-double, and quadruple-double precision floating-point arithmetic with AVX2

Acceleration of LU decomposition supporting double-double, triple-double, and quadruple-double precision floating-point arithmetic with AVX2
复制标题

使用 AVX2 加速 LU 分解,支持双双精度、三双精度和四双精度浮点运算

DOI:
10.1109/arith51176.2021.00021
复制
发表时间:
2021
期刊:
2021 IEEE 28th Symposium on Computer Arithmetic (ARITH)
影响因子:
--
通讯作者:
Kouya Tomonori
Kouya Tomonori
中科院分区:
--
文献类型:
--
作者:
Kouya Tomonori

文献摘要

被引文献

相似文献

在本文中,我们报告了用Intel的高级向量扩展2(AVX2)加速多二进制64型多精度LU分解的结果。我们的目标是使用某些类型的无误差变换(EFT)算法设计的双倍(DD)、三倍(TD)和四倍(QD)精度算法。在x86_计算环境下,利用SIMD化的EFT函数实现了加速的DD、TD和QD精度加法和乘法运算,在x86_计算环境下实现了四个二进制64位数的同时计算,从而实现了基于SIMD化矩阵乘法的多精度LU分解。我们的逻辑单元分解比未加速的快三倍。
In this paper, we report the results obtained from the acceleration of multi-binary64-type multiple precision LU decomposition with Intel's Advanced Vector Extensions 2 (AVX2). We targeted double-double (DD), triple-double (TD), and quad-double (QD) precision arithmetic designed using certain types of error-free transformation (EFT) arithmetic. We implemented accelerated DD, TD, and QD precision addition and multiplication using SIMDized EFT functions with AVX2, which perform simultaneous computation using four binary64 numbers on the x86_64 computing environment, and with these, we were able to develop multiple-precision LU decomposition based on SIMDized matrix multiplication. Our LU decomposition is up to three times faster than the non-accelerated one.