A high throughput 16 by 16 bit multiplier for DSP cores

A high throughput 16 by 16 bit multiplier for DSP cores
复制标题

用于 DSP 内核的高吞吐量 16 x 16 位乘法器

DOI:
10.1109/iscas.1996.541750
复制
发表时间:
1996
期刊:
1996 IEEE International Symposium on Circuits and Systems. Circuits and Systems Connecting the World. ISCAS 96
影响因子:
--
通讯作者:
C. Lemonds
C. Lemonds
中科院分区:
--
文献类型:
--
作者:
C. Lemonds

文献摘要

被引文献

相似文献

在快速发展的便携式电信行业中,数字信号处理 (DSP) 芯片是人们关注的焦点。对这些芯片更高性能的需求不断升级。随着对更高性能的需求,人们开始重视降低便携式电子产品的电源。降低电源电压会对性能产生负面影响。要在不降低 DSP 芯片性能的情况下降低功耗,需要重新评估架构、算法、电路设计和技术。电路风格是实现这些目标的关键因素。一段时间以来,静态 CMOS 因其稳健性和良好的性能而成为首选的传统技术。随着电源电压降低和性能需求增加,需要更快的电路风格,例如动态逻辑。双轨多米诺骨牌是一种可以实现这些目标的动态逻辑。乘法器在 DSP 内核中执行基本功能,并且可以根据应用大量使用。根据操作数的大小,乘法器可能占据芯片面积的很大一部分。相比之下,FPU 中的乘法器通常具有较大的操作数,从而产生较长的延迟路径。因此,它们通常是流水线化的,以便满足性能目标。本文重点介绍一个 16 x 16 阵列乘法器,其工作时钟频率为 500 MHz,具有四个周期的延迟。该乘法器采用双轨多米诺逻辑设计,采用 0.35 /spl mu/m CMOS 工艺。
In the rapidly growing portable telecommunications industry digital signal processing (DSP) chips are at the center of interest. The demand for higher performance for these chips continues to escalate. Coupled with demand for higher performance is the emphasis on lowering the power supply for portable electronics. Lowering the power supply has a negative impact on performance. Lowering the power supply with no degradation in performance for DSP chips requires a re-evaluation of architecture, algorithm, circuit design, and technology. Circuit style is a key factor in meeting these goals. Static CMOS has been the conventional technology of choice for some time because of its robustness and good performance. A faster circuit style such as dynamic logic is needed as power supplies go lower and performance needs increase. Dual rail domino is one type of dynamic logic that can meet these goals. Multipliers perform an essential function in DSP cores and can be heavily used depending upon the application. Depending upon the size of the operands, multipliers can occupy a significant portion of the chip area. In contrast multipliers in FPUs typically have large operands thus creating long delay paths. Hence, they are usually pipelined in order to meet performance goals. This paper focuses on a 16 by 16 array multiplier that operates at a clock frequency of 500 MHz with four cycles of latency. The multiplier is designed in dual rail domino logic using a 0.35 /spl mu/m CMOS process.