ReMCA: A Reconfigurable Multi-Core Architecture for Full RNS Variant of BFV Homomorphic Evaluation

ReMCA: A Reconfigurable Multi-Core Architecture for Full RNS Variant of BFV Homomorphic Evaluation
复制标题

ReMCA:用于 BFV 同态评估的完整 RNS 变体的可重构多核架构

DOI:
--
复制
发表时间:
2022
期刊:
IEEE Transactions on Circuits and Systems Part 1: Regular Papers
影响因子:
--
通讯作者:
Song Zhao
Song Zhao
中科院分区:
--
文献类型:
--
作者:
Yang Su;Bailong Yang;Chen Yang;Song Zhao

文献摘要

被引文献

相似文献

完全同态加密(FHE)允许对加密数据进行任意计算,因此在隐私保护计算方面具有潜力。然而,效率仍然是瓶颈。在本文中,我们提出了一种面积高效且高度统一的可重构多核架构(称为 ReMCA),用于 Brakerski 方案的 Fan-Vercauteren 变体(RNS-BFV)的完整残差数系统(RNS)变体,该架构采用可变数量的可重构处理元件(PE)和 RNS 通道。 PE单元可以灵活配置为NTT、INTT或模乘法器,从而避免了其他额外计算单元的需要。为了降低计算复杂度,ReMCA将前/后处理合并到NTT/INTT中,并统一NTT和INTT的读/写结构。此外,还提出了一种不需要单独的位反转操作的无冲突内存访问模式来优化内存访问。此外,针对不同的计算需求,引入了统一的硬件架构映射模型和数据内存组织模型,并将RNS-BFV涉及的所有计算单元都在ReMCA上进行了优化和映射。 ReMCA 在 Xilinx Virtex-7 FPGA 平台上进行评估。运行频率为250MHz,每秒可执行2260次同态乘法。当归一化为相同的参数集时,ReMCA 的吞吐量和区域时间乘积 (ATP) 实现了 $1.45 imes sim 5.51 imes $ 和 $1.58 imes sim 5.12 imes $ 的改进。
Fully homomorphic encryption (FHE) allows arbitrary computation on encrypted data and thus has potential in privacy-preserving computing. However, efficiency is still the bottleneck. In this paper we present an area-efficient and highly unified reconfigurable multi-core architecture (named ReMCA) for full Residue Number System (RNS) variant of Fan-Vercauteren variant of Brakerski’s scheme (RNS-BFV), which employs a variable number of reconfigurable processing elements (PEs) and RNS channels. The PE unit can be flexibly configured as NTT, INTT or modular multiplier, thereby avoiding the need of other extra computational units. To reduce the computational complexity, ReMCA merges the pre/post-processing into NTT/INTT and unifies the read/write structure of NTT and INTT. Also, a conflict-free memory access pattern that doesn’t need separate bit-reversal operation is proposed to optimize the memory access. Furthermore, targeting different computational requirements, a unified hardware architecture mapping model and data memory organization model are introduced, and all the computing units that RNS-BFV involved are optimized and mapped on ReMCA. ReMCA is evaluated on a Xilinx Virtex-7 FPGA platform. Running at 250MHz, it can perform 2260 homomorphic multiplication per second. When normalized to the same parameter set, the throughput and Area-Time-Products (ATPs) of ReMCA achieve $1.45 imes sim 5.51 imes $ and $1.58 imes sim 5.12 imes $ improvements.