A digital multiprocessor hardware accelerator board for cellular neural networks: CNN-HAC

A digital multiprocessor hardware accelerator board for cellular neural networks: CNN-HAC
复制标题

DOI:
10.1002/cta.4490200512
复制
发表时间:
1992-09
期刊:
Int. J. Circuit Theory Appl.
影响因子:
--
通讯作者:
T. Roska;G. Bártfai;P. Szolgay;T. Szirányi;A. Radványi;T. Kozek;Z. Ugray;Á. Zarándy
T. Roska;G. Bártfai;P. Szolgay;T. Szirányi;A. Radványi;T. Kozek;Z. Ugray;Á. Zarándy
中科院分区:
其他
文献类型:
--
作者:
T. Roska;G. Bártfai;P. Szolgay;T. Szirányi;A. Radványi;T. Kozek;Z. Ugray;Á. Zarándy

文献摘要

被引文献

相似文献

神经网络的并行实现在速度上具有上级优势。使用目录可编程VLSI IC的硬件加速器板代表了具有更高可重构性和更低成本的折衷。本文提出了一种细胞神经网络(CNN)的解决方案。给出了本设计(CNN-HAC)的结构,该结构使用四个标准DSP来计算包含(0.25-0.75)× 106个模拟神经细胞(取决于模板的类型)的单层CNN的瞬态响应。体系结构以及设计原理与处理器的数量无关。实际的设计是以PC附加板的形式进行的。全局控制单元主要用EPLD实现,它将板连接到主机固件,并将控制信号传递到DSP的本地控制单元。详细讨论了虚拟处理元件-计算模拟神经细胞的时间离散模型-与物理元件之间的特殊对应关系。它是在一个架构中实现的,具有简单的双向处理器间通信。这种架构可以使用更快的处理器、EPLD和存储器来“缩小”。当前版本以2 μs/cell/迭代速度运行。
Analogue realizations of neural networks are superior in speed. the hardware accelerator boards using catalogue programmable VLSI ICs represent a trade-off having higher reconfigurability and lower cost. This paper presents such a solution for a cellular neural network (CNN). The architecture of the present design (CNN-HAC) using four standard DSPs to calculate the transient response of a one-layer CNN containing (0.25–0.75) × 106 analogue neural cells (depending on the type of template) is presented. the architecture and also the design principles are independent of the number of processors. the actual design was made in the form of a PC add-on board. The global control unit, which connects the board to the host firmware and communicates control signals to/from the local control units of the DSPs, was realized mainly with EPLDs. A special correspondence between the virtual processing elements—calculating the time-discrete models of the analogue neural cells—and the physical ones is discussed in detail. It is realized in an architecture with a simple, two-directional interprocessor communication. This architecture can be ‘scaled down’ using faster processors, EPLDs and memories. the present version runs with 2 μs/cell/iteration speed.