A Tag Based Random Order Vector Reduction Circuit

A Tag Based Random Order Vector Reduction Circuit
复制标题

基于标签的随机顺序向量约简电路

DOI:
10.1109/access.2020.2976764
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Wei Ming
Wei Ming
中科院分区:
计算机科学3区
文献类型:
--
作者:
Huang Yihua;Huang Wenjin;Chen Rui;Wu Huangtao;Wei Ming

文献摘要

参考文献

相似文献

向量归约是在许多科学和工程应用场景中将向量归约为单个标量值的非常常见的操作。因此,快速高效的矢量归约电路对实时系统的应用具有重要意义。通常广泛采用流水线结构来提高矢量归约电路的吞吐量并实现最大效率。针对随机输入序列中多个变长向量的处理问题,提出了一种基于标签的全流水线向量归约电路,该电路通过缓存状态模块查询和更新每个向量的该高速缓存状态。然而,当输入向量的数量变大时,需要更大的高速缓存状态模块,这消耗更多的组合逻辑并且降低操作频率。针对这一问题,提出了一种高速缓存电路,将输入矢量分成若干组,送入专用的缓存状态电路,提高了工作频率。与已有的工作相比,该原型电路和基于原型电路的改进电路在不同输入矢量长度下都能获得最小的Slices $ {\times }$ us(小于现有工作的80%)。此外,这两种电路都能提供简单而有效的接口,其存取时序与RAM的存取时序相似。因此,该电路可以在更大的范围内应用。
Vector reduction is a very common operation to reduce a vector into a single scalar value in many scientific and engineering application scenarios. Therefore a fast and efficient vector reduction circuit has great significance to the real-time system applications. Usually the pipeline structure is widely adopted to increase the throughput of the vector reduction circuit and achieve maximum efficiency. In this paper, to deal with multiple vectors of variable length in random input sequence, a novel tag based fully pipelined vector reduction circuit is firstly proposed, in which a cache state module is used to queer and update the cache state of each vector. However, when the quantity of the input vector becomes large, a larger cache state module is required, which consumes more combinational logic and lower the operating frequency. To solve this problem, a high speed circuit is proposed in which the input vectors will be divided into several groups and sent into the dedicated cache state circuits, which can improve the operating frequency. Compared with other existing work, the prototype circuit and the improved circuit based on the prototype circuit can achieve the smallest Slices $ {\times }$ us (<80% of the state-of-the-art work) for different input vector lengths. Moreover, both circuits can provide simple and efficient interface whose access timing is similar to that of a RAM. Therefore the circuits can be applied in a greater range.
DOI: 10.1109/12.841125
发表时间: 2000-03
期刊: IEEE Trans. Computers
影响因子: --
作者:
Zhen Luo;M. Martonosi
通讯作者: Zhen Luo;M. Martonosi
DOI: --
发表时间: 1981
期刊: --
影响因子: --
作者:
P. Kogge
通讯作者: P. Kogge
DOI: 10.1109/tc.1985.1676580
发表时间: 1985-05
影响因子: 3.7
作者:
L. Ni;K. Hwang
通讯作者: L. Ni;K. Hwang
DOI: 10.1109/tpds.2011.141
发表时间: 2012-02
影响因子: 5.3
作者:
Yi-Gang Tai;D. Lo;K. Psarris
通讯作者: Yi-Gang Tai;D. Lo;K. Psarris
DOI: 10.1109/12.73591
发表时间: 1991-02
期刊: IEEE Trans. Computers
影响因子: --
作者:
H. Sips;H. Lin
通讯作者: H. Sips;H. Lin