High-Speed and Low-Latency ECC Processor Implementation Over GF(2m) on FPGA

High-Speed and Low-Latency ECC Processor Implementation Over GF(2m) on FPGA
复制标题

DOI:
10.1109/tvlsi.2016.2574620
复制
发表时间:
2017-01-01
影响因子:
2.8
通讯作者:
Benaissa, Mohammed
Benaissa, Mohammed
中科院分区:
工程技术2区
文献类型:
--
作者:
Khan, Zia U. A.;Benaissa, Mohammed

文献摘要

被引文献

相似文献

提出了一种基于现场可编程门阵列(FPGA)的高速椭圆曲线密码(ECC)点乘(PM)处理器实现方案。一个新的分段流水线的全精度乘法器被用来减少延迟,和Lopez-Dahab蒙哥马利PM算法被修改为仔细调度,以避免数据依赖性,从而导致在所需的时钟周期(CC)的数量急剧减少。所提出的ECC架构已在Xilinx FPGA的Virtex 4、Virtex 5和Virtex 7系列上实现。据我们所知,我们的单和三个基于乘法器的设计显示出最快的性能,到目前为止,与个别报道的作品相比。我们的基于一个乘法器的ECC处理器在Virtex 4(210 MHz时为5.32 mu s)、Virtex 5(228 MHz时为4.91 mu s)和更高级的Virtex 7(352 MHz时为3.18 mu s)上也实现了最高的报告速度和最佳的报告面积-时间性能。最后,建议三乘法器为基础的ECC实现是第一个工作报告的CC数量最少,最快的ECC处理器设计的FPGA(450 CC,以获得2.83亩Virtex 7)。
In this paper, a novel high-speed elliptic curve cryptography (ECC) processor implementation for point multiplication (PM) on field-programmable gate array (FPGA) is proposed. A new segmented pipelined full-precision multiplier is used to reduce the latency, and the Lopez-Dahab Montgomery PM algorithm is modified for careful scheduling to avoid data dependency resulting in a drastic reduction in the number of clock cycles (CCs) required. The proposed ECC architecture has been implemented on Xilinx FPGAs' Virtex4, Virtex5, and Virtex7 families. To the best of our knowledge, our single-and three-multiplier-based designs show the fastest performance to date when compared with reported works individually. Our one-multiplier-based ECC processor also achieves the highest reported speed together with the best reported area-time performance on Virtex4 (5.32 mu s at 210 MHz), on Virtex5 (4.91 mu s at 228 MHz), and on the more advanced Virtex7 (3.18 mu s at 352 MHz). Finally, the proposed three-multiplier-based ECC implementation is the first work reporting the lowest number of CCs and the fastest ECC processor design on FPGA (450 CCs to get 2.83 mu s on Virtex7).