NTTGen: a framework for generating low latency NTT implementations on FPGA

NTTGen: a framework for generating low latency NTT implementations on FPGA
复制标题

DOI:
10.1145/3528416.3530225
复制
发表时间:
2022-05
期刊:
Proceedings of the 19th ACM International Conference on Computing Frontiers
影响因子:
--
通讯作者:
Yang Yang-Yang;S. Kuppannagari;R. Kannan;V. Prasanna
Yang Yang-Yang;S. Kuppannagari;R. Kannan;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Yang Yang-Yang;S. Kuppannagari;R. Kannan;V. Prasanna

文献摘要

被引文献

相似文献

同态加密(HE)是一种有前途的技术,可确保云中应用程序的安全和隐私。数论变换 (NTT) 是基于 HE 的应用程序中的关键操作。 HE 需要截然不同的 NTT 参数来满足应用程序的性能和安全要求。 FPGA 不断增强的计算能力和灵活性使其对于加速 NTT 具有吸引力。然而,FPGA 编程仍然需要硬件设计专业知识和大量的开发工作。为了缩小差距,我们提出了 NTTGen,这是一个自动生成针对基于 HE 的应用程序的低延迟 NTT 设计的框架。 NTTGen 将应用参数、延迟和硬件资源约束作为输入,确定设计参数,并生成可综合的 Verilog 代码作为输出。低延迟 NTT 实现是通过改变数据、管道和批处理并行性来实现的。 NTTGen 利用流排列网络来降低 NTT 计算中各阶段之间的互连复杂性。该框架支持两种类型的 NTT 内核来执行模运算(NTT 中的关键计算):用于特定类素数模的低延迟且资源高效的 NTT 内核和用于其他素数的通用 NTT 内核。我们进一步开发了设计空间探索流程,以确定最佳设计的硬件设计参数。我们通过生成各种 NTT 参数的设计来评估 NTTGen。与最先进的 FPGA 实现相比,该设计的延迟时间提高了 2.9 倍。
Homomorphic encryption (HE) is a promising technique to ensure the security and privacy of applications in the cloud. Number Theoretic Transform (NTT) is a key operation in HE-based applications. HE requires vastly different NTT parameters to meet the performance and security requirements of applications. The increasing compute capabilities and flexibility of FPGAs make them attractive to accelerate NTT. However, programming FPGA still involves hardware design expertise and significant development effort. To close the gap, we propose NTTGen, a framework to automatically generate low latency NTT designs targeting HE-based applications. NTTGen takes application parameters, latency and hardware resource constraints as input, determines the design parameters, and produces synthesizable Verilog code as output. Low latency NTT implementations are obtained by varying the data, pipeline and batch parallelism. NTTGen utilizes streaming permutation network to reduce the interconnect complexity between stages in the NTT computation. The framework supports two types of NTT cores to perform modular arithmetic, the key computation in NTT: a low latency and resource efficient NTT core for a specific class of prime moduli and a general purpose NTT core for other primes. We further develop a design space exploration flow to identify the hardware design parameters of an optimal design. We evaluate NTTGen by generating designs for various NTT parameters. The designs result in up to 2.9X improvement in latency over the state-of-the-art FPGA implementations.