Low-Latency, Low-Area, and Scalable Systolic-Like Modular Multipliers for GF(2m) Based on Irreducible All-One Polynomials

Low-Latency, Low-Area, and Scalable Systolic-Like Modular Multipliers for GF(2m) Based on Irreducible All-One Polynomials
复制标题

DOI:
10.1109/tcsi.2016.2614309
复制
发表时间:
2017-02-01
影响因子:
5.1
通讯作者:
Lou, Xin
Lou, Xin
中科院分区:
工程技术2区
文献类型:
--
作者:
Meher, Pramod Kumar;Lou, Xin

文献摘要

被引文献

相似文献

本文提出了一种基于不可约AOP的GF(2(m))上典型基有限域乘法的脉动实现的高效递归公式。我们已经推导出一个递归算法的乘法,并使用它来设计一个定期和本地化的位级依赖图(DG)的脉动计算。通过节点分裂将位级规则DG转换为细粒度DG,并将其映射到并行脉动结构中。与大多数现有结构不同的是,它不涉及任何全球性的模块化缩减通信。所提出的位并行脉动结构具有与现有的最佳位并行脉动结构相同的周期时间[1],但涉及的寄存器数量明显较少。所提出的位并行设计具有l + [ log(2)s]+1个周期的可缩放延迟,这与现有的脉动设计相比相当低。此外,所提出的时分复用结构是专门为吞吐量和硬件复杂性的可扩展性而设计的,以满足资源受限应用中的面积-时间权衡,同时保持或减少整体延迟。ASIC综合报告表明,所提出的位并行结构提供了近30%的节省面积和近38%的节省功耗比现有的最好的基于AOP的脉动有限域乘法器。
In this paper, an efficient recursive formulation is suggested for systolic implementation of canonical basis finite field multiplication over GF(2(m)) based on irreducible AOP. We have derived a recursive algorithm for the multiplication, and used that to design a regular and localized bit-level dependence graph (DG) for systolic computation. The bit-level regular DG is converted into a fine-grained DG by node-splitting, and mapped that into a parallel systolic architecture. Unlike most of the existing structures, it does not involve any global communications for modular reduction. The proposed bit-parallel systolic structure has the same cycle time as that of the best existing bit-parallel systolic structure [1], but involves significantly less number of registers. The proposed bit-parallel design has a scalable latency of l + [ log(2)s] + 1 cycles which is considerably low compared with those of existing systolic designs. Moreover, the proposed time-multiplexed structure is designed specifically for scalability of throughput and hardware-complexity to meet the area-time trade-off in resource-constrained applications while maintaining or reducing the overall latency. The ASIC synthesis report shows that the proposed bit-parallel structures offers nearly 30% saving of area and nearly 38% saving of power consumption over the best of the existing AOP-based systolic finite field multiplier.