Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures

Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures
复制标题

DOI:
10.1109/tip.2017.2678799
复制
发表时间:
2017-05
影响因子:
10.6
通讯作者:
Cesar Carranza;D. Llamocca;M. Pattichis
Cesar Carranza;D. Llamocca;M. Pattichis
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cesar Carranza;D. Llamocca;M. Pattichis

文献摘要

被引文献

相似文献

该手稿描述了快速和可扩展的架构和相关算法计算卷积和互相关。基本思想是将2D卷积和互相关映射到变换域中的1D卷积和互相关的集合。这是通过对一般核使用离散周期Radon变换和对低秩核使用奇异值分解-LU分解来实现的。该方法使用可扩展的架构,可以适应现代FPGA和Zynq-SOC设备。基于不同类型的可用资源,对于$P\times P$块,2D卷积和互相关可以在$O(P)$时钟周期内计算,最多可达$O(P^{2})$时钟周期。因此,在性能与所需资源的数量和类型之间存在权衡。我们使用现代可编程器件(Virtex-7和Zynq-SOC)提供所提出的架构的实现。基于所需资源的数量和类型,我们表明,所提出的方法显着优于目前的方法。
The manuscript describes fast and scalable architectures and associated algorithms for computing convolutions and cross-correlations. The basic idea is to map 2D convolutions and cross-correlations to a collection of 1D convolutions and cross-correlations in the transform domain. This is accomplished through the use of the discrete periodic radon transform for general kernels and the use of singular value decomposition -LU decompositions for low-rank kernels. The approach uses scalable architectures that can be fitted into modern FPGA and Zynq-SOC devices. Based on different types of available resources, for $P\times P$ blocks, 2D convolutions and cross-correlations can be computed in just $O(P)$ clock cycles up to $O(P^{2})$ clock cycles. Thus, there is a trade-off between performance and required numbers and types of resources. We provide implementations of the proposed architectures using modern programmable devices (Virtex-7 and Zynq-SOC). Based on the amounts and types of required resources, we show that the proposed approaches significantly outperform current methods.