Accelerating Number Theoretic Transformations for Bootstrappable Homomorphic Encryption on GPUs

Accelerating Number Theoretic Transformations for Bootstrappable Homomorphic Encryption on GPUs
复制标题

加速 GPU 上可引导同态加密的数论转换

DOI:
--
复制
发表时间:
2020
期刊:
IEEE International Symposium on Workload Characterization
影响因子:
--
通讯作者:
Jung Ho Ahn
Jung Ho Ahn
中科院分区:
--
文献类型:
--
作者:
Sangpyo Kim;Wonkyung Jung;J. Park;Jung Ho Ahn

文献摘要

被引文献

相似文献

同态加密(HE)引起了人们的极大关注,因为它提供了一种对加密消息进行隐私保护计算的方法。数论变换(NTT)是有限整数域上离散傅立叶变换(DFT)的一种特殊形式,是实现HE加密密文快速计算的关键算法。之前的工作通过利用DFT优化技术在流行的并行处理平台GPU上加速了NTT及其逆变换。然而,这些基于GPU的研究缺乏对NTT和DFT之间的主要差异的全面分析,或者仅考虑小的HE参数,这些HE参数在不解密的情况下可以执行的算术运算的数量方面具有严格的约束。在本文中,我们分析了NTT和DFT的算法特性,并评估NTT的性能时,我们通常适用于DFT和NTT在现代GPU上的优化。从分析中,我们确定,NTT遭受严重的主存带宽瓶颈的大型HE参数集。为了解决主存带宽的问题,我们提出了一种新的NTT特定的飞行根生成计划被称为飞行旋转(OT)。与基线基2 NTT实现相比,在应用所有优化(包括OT)后,我们在现代GPU上实现了4.2倍的加速比。
Homomorphic encryption (HE) draws huge attention as it provides a way of privacy-preserving computations on encrypted messages. Number Theoretic Transform (NTT), a specialized form of Discrete Fourier Transform (DFT) in the finite field of integers, is the key algorithm that enables fast computation on encrypted ciphertexts in HE. Prior works have accelerated NTT and its inverse transformation on a popular parallel processing platform, GPU, by leveraging DFT optimization techniques. However, these GPU-based studies lack a comprehensive analysis of the primary differences between NTT and DFT or only consider small HE parameters that have tight constraints in the number of arithmetic operations that can be performed without decryption. In this paper, we analyze the algorithmic characteristics of NTT and DFT and assess the performance of NTT when we apply the optimizations that are commonly applicable to both DFT and NTT on modern GPUs. From the analysis, we identify that NTT suffers from severe main-memory bandwidth bottleneck on large HE parameter sets. To tackle the main-memory bandwidth issue, we propose a novel NTT-specific on-the-fly root generation scheme dubbed on-the-fly twiddling (OT). Compared to the baseline radix-2 NTT implementation, after applying all the optimizations, including OT, we achieve 4.2⨯ speedup on a modern GPU.