A High-Performance Domain-Specific Processor With Matrix Extension of RISC-V for Module-LWE Applications

A High-Performance Domain-Specific Processor With Matrix Extension of RISC-V for Module-LWE Applications
复制标题

具有 RISC-V 矩阵扩展功能的高性能特定领域处理器,适用于模块 LWE 应用

DOI:
--
复制
发表时间:
2022
期刊:
IEEE Transactions on Circuits and Systems Part 1: Regular Papers
影响因子:
--
通讯作者:
Jun Han
Jun Han
中科院分区:
--
文献类型:
--
作者:
Yifan Zhao;Ruiqi Xie;Guozhu Xin;Jun Han

文献摘要

被引文献

相似文献

5G边缘计算基础设施应通过实施后量子密码学(PQC)来增强抗量子攻击能力。在各种PQC方案中,基于错误学习(LWE)的基于格的密码学(LBC)因其性能效率和安全保证而备受关注。在基于 LWE 的 LBC 中,基于模块 LWE 的方案受益于独特的多项式矩阵和向量结构,比其他方案更具优势。为了为边缘计算范例提供 Module-LWE 应用程序的高性能实现,我们提出了一种基于 RISC-V 架构矩阵扩展的特定领域处理器。此自定义扩展使用高级功能抽象封装了基于矩阵的环运算。提出了一种具有可配置功能的二维脉动阵列来执行基于矩阵的数论变换(NTT)和其他算术运算,在支持可变大小的多项式矩阵和向量结构的情况下实现高数据级并行性。由于Module-LWE的这种结构不涉及不同内部元素之间的数据依赖性,因此进一步开发了乱序机制来利用指令级并行性。我们在 TSMC 28nm 技术下实现了所提出的架构。评估结果表明,与最先进的加密处理器同行相比,我们的实施在 Kyber 和 Dilithium 中分别可以实现高达 3.5 美元和 3.3 美元的周期计数改进。
The 5G edge computing infrastructure should be empowered with quantum attack resistance by implementing post-quantum cryptography (PQC). Among various PQC schemes, lattice-based cryptography (LBC) based on learning with error (LWE) has attracted much attention because of its performance efficiency and security guarantee. In LWE-based LBCs, the Module-LWE-based schemes gain advantage over the others benefiting from the unique polynomial matrix and vector structure. To provide a high-performance implementation of Module-LWE applications for the edge computing paradigm, we propose a domain-specific processor based on a matrix extension of RISC-V architecture. This custom extension encapsulates the matrix-based ring operations with a high-level functional abstraction. A 2-D systolic array with configurable functionality is proposed to perform matrix-based number theoretic transform (NTT) and other arithmetic operations, achieving high data-level parallelism with support for the variable-sized polynomial matrix and vector structure. As this structure of Module-LWE involves no data dependency between different inner elements, an out-of-order mechanism is further developed to exploit the instruction-level parallelism. We implement the proposed architecture under TSMC 28nm technology. The evaluation results show that our implementation can achieve up to $3.5 imes $ and $3.3 imes $ improvement in cycle count respectively in Kyber and Dilithium, compared to the state-of-the-art crypto-processor counterparts.