Efficient Hardware Arithmetic for Inverted Binary Ring-LWE Based Post-Quantum Cryptography

Efficient Hardware Arithmetic for Inverted Binary Ring-LWE Based Post-Quantum Cryptography
复制标题

DOI:
10.1109/tcsi.2022.3169471
复制
发表时间:
2022-08
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
J. Imaña;Pengzhou He;Tianyou Bao;Yazheng Tu;Jiafeng Xie
J. Imaña;Pengzhou He;Tianyou Bao;Yazheng Tu;Jiafeng Xie
中科院分区:
其他
文献类型:
--
作者:
J. Imaña;Pengzhou He;Tianyou Bao;Yazheng Tu;Jiafeng Xie

文献摘要

被引文献

相似文献

基于Ring learning-with-errors(RLWE)的加密方案是一种基于格的密码算法,由于其高效的实现和低计算复杂度,其构成了后量子密码学(PQC)标准化的最有前途的候选者之一。二进制环LWE(BRLWE)是RLWE的一种新的优化变体,它实现了更小的计算复杂度和更高的硬件实现效率。本文提出了两种基于线性反馈移位寄存器(LFSR)的高效结构,用于基于倒置二进制环LWE(InvBRLWE)的加密算法,即多项式环$\mathbb {Z} q}/(x^{n}+1)$上的$A\cdot B+C$运算。第一种架构优化了主要计算的资源使用,并具有新颖的输入处理设置,以最小化输入加载周期来加速整体处理延迟。第二个架构部署了一个创新的串行输入串行输出处理格式,以减少所涉及的面积使用进一步保持定期输入加载的时间复杂度。实验结果表明,这里提出的架构改善了文献中发现的竞争方案所获得的复杂性,例如,涉及比最近的设计少71.23%的面积延迟产品。这两种架构在面积-时间复杂度方面都非常高效,并且可以扩展到不同的轻量级应用程序环境中。
Ring learning-with-errors (RLWE)-based encryption scheme is a lattice-based cryptographic algorithm that constitutes one of the most promising candidates for Post-Quantum Cryptography (PQC) standardization due to its efficient implementation and low computational complexity. Binary Ring-LWE (BRLWE) is a new optimized variant of RLWE, which achieves smaller computational complexity and higher efficient hardware implementations. In this paper, two efficient architectures based on Linear-Feedback Shift Register (LFSR) for the arithmetic used in Inverted Binary Ring-LWE (InvBRLWE)-based encryption scheme are presented, namely the operation of $A\cdot B+C$ over the polynomial ring $\mathbb {Z}_{q}/(x^{n}+1)$ . The first architecture optimizes the resource usage for major computation and has a novel input processing setup to speed up the overall processing latency with minimized input loading cycles. The second architecture deploys an innovative serial-in serial-out processing format to reduce the involved area usage further yet maintains a regular input loading time-complexity. Experimental results show that the architectures presented here improve the complexities obtained by competing schemes found in the literature, e.g., involving 71.23% less area-delay product than recent designs. Both architectures are highly efficient in terms of area-time complexities and can be extended for deploying in different lightweight application environments.