Highly-Parallel Hardwired Deep Convolutional Neural Network for 1-ms Dual-Hand Tracking

Highly-Parallel Hardwired Deep Convolutional Neural Network for 1-ms Dual-Hand Tracking
复制标题

DOI:
10.1109/tcsvt.2021.3103784
复制
发表时间:
2022-12
影响因子:
8.4
通讯作者:
Peiqi Zhang;Tingting Hu;Dingli Luo;Songlin Du;T. Ikenaga
Peiqi Zhang;Tingting Hu;Dingli Luo;Songlin Du;T. Ikenaga
中科院分区:
工程技术1区
文献类型:
--
作者:
Peiqi Zhang;Tingting Hu;Dingli Luo;Songlin Du;T. Ikenaga

文献摘要

相似文献

1-ms视觉系统代表视频感测技术中时间发展的极端情况。此外,1毫秒的双手跟踪系统利用了手的灵巧功能,因此可以作为人机交互的无缝和直观的界面。深度CNN有望实现高跟踪鲁棒性,然而,基于GPU和基于FPGA的实现都无法以超高速解决跟踪任务。本文提出:(a)直接将深度CNN映射为硬连线电路的范例,因此整个网络并行运行并获得高处理速度。网络被免除存储器访问,因为所有中间神经值都隐式地表示在硬件状态中。(B)在FPGA上实现了硬件网络的设计,在硬件网络中设计了核自适应卷积树,以最大化并行度。因此,通过将卷积层实现为具有统一组件的细粒度管道,消除了网络的速度瓶颈;(c)FPGA-GPU异质互补,利用辅助GPU网络来补偿FPGA网络的准确性,而不影响其速度。FPGA上的快速初步结果使用来自GPU的延迟但准确的提示进行间歇性改进。实验结果表明,该方法对640 × 480$的图像处理速度达到973 fps,处理时间仅为1.30 ms,而对测试序列的准确率仅比常规方法低4.7%。视频演示可在https://wcms.waseda.jp/em/5f9d020f136e7上获得。
1-ms vision systems represent an extreme case of temporal development in video sensing techniques. Moreover, a 1-ms dual-hand tracking system leverages the dexterous functionality of hands and thus serves as a seamless and intuitive interface for Human-Computer Interaction. Deep CNN is promising for high tracking robustness, however, neither GPU-based nor FPGA-based implementation addresses the tracking task with ultra-high-speed. This paper proposes: (a) A paradigm to directly map a deep CNN as a hardwired circuit, so the entire network runs in parallel and high processing speed is obtained. The network is exempted from memory access since all intermediate neural values are implicitly represented in hardware states. And condensed binarization is used to reduce resource utilization; (b) Hardware design of the hardwired network on FPGA, inside which kernel-adapted convolutional trees are devised to maximize the parallelism. The speed bottleneck of the network is therefore removed by implementing convolutional layers as fine-grained pipelines with unified components; (c) FPGA-GPU hetero complementation, which utilizes an auxiliary GPU network to compensate for accuracy of the FPGA network without affecting its speed. The quick primary results on FPGA are intermittently refined using delayed but accurate hints from GPU. Implementation results show that the proposed method reaches 973fps and consumes merely 1.30ms to process on $640\times 480$ images, while the accuracy is only 4.7% lower compared with the general method on test sequences. Video demonstrations are available at https://wcms.waseda.jp/em/5f9d020f136e7.