Parallel modular multiplication using 512-bit advanced vector instructions

Parallel modular multiplication using 512-bit advanced vector instructions
复制标题

使用 512 位高级向量指令的并行模乘法

DOI:
10.1007/s13389-021-00256-9
复制
发表时间:
2021
影响因子:
1.9
通讯作者:
C. Haider
C. Haider
中科院分区:
计算机科学4区
文献类型:
--
作者:
B. Buhrow;B. Gilbert;C. Haider

文献摘要

被引文献

相似文献

像公钥密码这样的应用程序的性能严重依赖于模乘的速度。本文介绍了一种新的基于块的蒙哥马利乘法的变体,块积扫描(BPS)方法,该方法在现代英特尔处理器系列上使用新的512位高级向量指令(AVX-512)时特别有效。我们的并行乘法方法还允许平方和次二次Karatsuba增强。我们在英特尔至强CPU上演示了与OpenSSL相比解密吞吐量的改进,以及与GMP-6.1.2相比模幂运算吞吐量的改进。此外,与众核Knights Landing Xeon Phi硬件上最先进的矢量实现相比,我们展示了解密吞吐量的改进。最后,我们展示了如何交错中国剩余定理为基础的RSA计算在我们的并行BPS技术一半的解密延迟,同时提供保护,防止故障注入攻击。
Applications such as public-key cryptography are critically reliant on the speed of modular multiplication for their performance. This paper introduces a new block-based variant of Montgomery multiplication, the Block Product Scanning (BPS) method, which is particularly efficient using new 512-bit advanced vector instructions (AVX-512) on modern Intel processor families. Our parallel-multiplication approach also allows for squaring and sub-quadratic Karatsuba enhancements. We demonstrateimprovement in decryption throughput in comparison with OpenSSL andimprovement in modular exponentiation throughput compared to GMP-6.1.2 on an Intel Xeon CPU. In addition, we showimprovement in decryption throughput in comparison with state-of-the-art vector implementations on many-core Knights Landing Xeon Phi hardware. Finally, we show how interleaving Chinese remainder theorem-based RSA calculations within our parallel BPS technique halves decryption latency while providing protection against fault-injection attacks.