A High-performance Hardware Implementation of Saber Based on Karatsuba Algorithm

A High-performance Hardware Implementation of Saber Based on Karatsuba Algorithm
复制标题

基于Karatsuba算法的Sabre高性能硬件实现

DOI:
--
复制
发表时间:
2020
期刊:
IACR Cryptology ePrint Archive
影响因子:
--
通讯作者:
Leibo Liu
Leibo Liu
中科院分区:
--
文献类型:
--
作者:
Yihong Zhu;Min Zhu;Bohan Yang;Wenping Zhu;Chenchen Deng;Chen Chen;Shaojun Wei;Leibo Liu

文献摘要

被引文献

相似文献

.抽象。虽然已经提出了大量的硬件和软件实现来加速基于格的密码学,Saber,一个基于模块LWR的算法,它已经推进到NIST标准化过程的第二轮,还没有得到足够的支持,目前的解决方案。基于这些动机,本文提出了一种基于算法-硬件协同设计的高性能密码处理器。首先,一个分层的Karatsuba计算框架,一个硬件高效的Karatsuba调度策略和优化的电路结构,以实现高吞吐量的多项式乘法。此外,提出了任务级流水线和截断乘法器,以实现算法特定的细粒度处理。通过上述所有优化,我们的处理器分别需要943、1156和408个时钟周期来进行密钥生成、加密和解密。通过这些优化,我们的处理器需要943,1156和408个时钟周期的密钥生成,加密和Saber 768解密,实现5.4倍,5.2倍和4.2倍减少相比,最先进的FPGA解决方案,分别。我们的设计的布局后模拟实现与台积电40纳米CMOS工艺在0.35平方毫米。Saber 768的吞吐量高达每秒346 k次加密操作,能量效率为0.12 uJ/加密,同时工作在400 MHz,与当前PQC硬件解决方案相比,分别实现了近52倍和30倍的改进。
. Abstract. Although large numbers of hardware and software implementations have been proposed to accelerate lattice-based cryptography, Saber, a module-LWR-based algorithm, which has advanced to second round of the NIST standardization process, has not been adequately supported by the current solutions. Based on these motivations, a high-performance crypto-processor is proposed based on an algorithm-hardware co-design in this paper. First, a hierarchical Karatsuba calculating framework, a hardware-efficient Karatsuba scheduling strategy and an optimized circuit structure are utilized to enable high-throughput polynomial multiplication. Furthermore, a task-level pipeline and truncated multipliers are proposed to enable algorithm-specific fine-grained processing. Enabled by all of the above optimizations, our processor takes 943, 1156, and 408 clock cycles for key generation, encryption, and de-cryption, respectively. Enabled by these optimizations, our processor takes 943, 1156 and 408 clock cycles for key generation, encryption, and decryption of Saber768, achieving 5.4 × , 5.2 × and 4.2 × reductions compared with the state-of-the-art FPGA solutions, respectively. The post-layout simulation of our design is implemented with TSMC 40 nm CMOS process within 0.35 mm 2 . The throughput for Saber768 is up to 346k encryption operations per second and the energy efficiency is 0.12 uJ/encryption while operating at 400 MHz, achieving nearly 52 × improvement and 30 × improvement, respectively compared with current PQC hardware solutions.