Exploring Parallelism to Improve the Performance of FrodoKEM in Hardware

Exploring Parallelism to Improve the Performance of FrodoKEM in Hardware
复制标题

DOI:
10.1007/s13389-021-00258-7
复制
发表时间:
2021-03
影响因子:
1.9
通讯作者:
James Howe;Marco Martinoli;E. Oswald;F. Regazzoni
James Howe;Marco Martinoli;E. Oswald;F. Regazzoni
中科院分区:
计算机科学4区
文献类型:
--
作者:
James Howe;Marco Martinoli;E. Oswald;F. Regazzoni

文献摘要

被引文献

相似文献

FrodoKEM是一种基于晶格的密钥封装机制,目前是NIST后量子标准化工作的半决赛入围者。这些候选者的一个条件是对随机性来源使用NIST标准(即种子扩展),因此大多数候选者使用SHA-3标准中定义的SHARK。然而,对于许多考生来说,这个模块是一个重大的实现瓶颈。Trivium是一种轻量级的ISO标准流密码,在硬件上表现良好,并已用于先前的基于格的密码学的硬件设计。这项研究提出了FrodoKEM的优化设计,通过将密码方案中的矩阵乘法运算并行化来专注于高吞吐量。由于Trivium具有更高的产量和更低的面积消耗,因此使用Trivium可以简化这一过程。建议的并行化也补充了解封模块中增加的一阶掩蔽。总体而言,我们显著提高了FrodoKEM的吞吐量;对于封装,我们看到了加速,达到了每秒825次操作;对于解封装,我们看到了加速,与以前的技术水平相比,达到了每秒763次操作,同时还保持了类似的不到2000个片的FPGA面积占用。
FrodoKEM is a lattice-based key encapsulation mechanism, currently a semi-finalist in NIST’s post-quantum standardisation effort. A condition for these candidates is to use NIST standards for sources of randomness (i.e. seed-expanding), and as such most candidates utilise SHAKE, an XOF defined in the SHA-3 standard. However, for many of the candidates, this module is a significant implementation bottleneck. Trivium is a lightweight, ISO standard stream cipher which performs well in hardware and has been used in previous hardware designs for lattice-based cryptography. This research proposes optimised designs for FrodoKEM, concentrating on high throughput by parallelising the matrix multiplication operations within the cryptographic scheme. This process is eased by the use of Trivium due to its higher throughput and lower area consumption. The parallelisations proposed also complement the addition of first-order masking to the decapsulation module. Overall, we significantly increase the throughput of FrodoKEM; for encapsulation we see aspeed-up, achieving 825 operations per second, and for decapsulation we see aspeed-up, achieving 763 operations per second, compared to the previous state of the art, whilst also maintaining a similar FPGA area footprint of less than 2000 slices.