Acceleration of LU decomposition supporting double-double, triple-double, and quadruple-double precision floating-point arithmetic with AVX2
Acceleration of LU decomposition supporting double-double, triple-double, and quadruple-double precision floating-point arithmetic with AVX2
复制标题
使用 AVX2 加速 LU 分解,支持双双精度、三双精度和四双精度浮点运算
DOI:
10.1109/arith51176.2021.00021
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Kouya Tomonori
中科院分区:
文献类型:
--
作者:
Kouya Tomonori
In this paper, we report the results obtained from the acceleration of multi-binary64-type multiple precision LU decomposition with Intel's Advanced Vector Extensions 2 (AVX2). We targeted double-double (DD), triple-double (TD), and quad-double (QD) precision arithmetic designed using certain types of error-free transformation (EFT) arithmetic. We implemented accelerated DD, TD, and QD precision addition and multiplication using SIMDized EFT functions with AVX2, which perform simultaneous computation using four binary64 numbers on the x86_64 computing environment, and with these, we were able to develop multiple-precision LU decomposition based on SIMDized matrix multiplication. Our LU decomposition is up to three times faster than the non-accelerated one.