RoadNet-RT: High Throughput CNN Architecture and SoC Design for Real-Time Road Segmentation

RoadNet-RT: High Throughput CNN Architecture and SoC Design for Real-Time Road Segmentation
复制标题

DOI:
10.1109/tcsi.2020.3038139
复制
发表时间:
2021-02-01
影响因子:
5.1
通讯作者:
Huang, Xinming
Huang, Xinming
中科院分区:
工程技术2区
文献类型:
--
作者:
Bai, Lin;Lyu, Yecheng;Huang, Xinming

文献摘要

被引文献

相似文献

近年来,卷积神经网络(CNN)在许多工程应用中越来越受欢迎,特别是在计算机视觉方面。为了获得更好的性能,神经网络中融入了更复杂的结构和高级操作,这导致推理时间非常长。对于自动驾驶和虚拟现实等时间关键型任务,实时处理是基础。为了达到实时处理速度,本文提出了一种轻量级,高吞吐量的CNN架构,即RoadNet-RT,用于道路分割。该算法在KITTI道路分割数据集上获得了92.55%的MaxF分数。在GTX 1080 GPU上运行时,推理时间约为每帧9 ms。与最先进的网络相比,RoadNet-RT将推理时间加快了17.8倍,而精度损失仅为3.75%。在CamVid数据集上,其准确率为92.98%。在硬件加速器的设计中,对深度可分离卷积、非均匀核卷积等技术进行了优化。所提出的CNN架构已成功地实现在ZCU 102 MPSoC FPGA上,使用INT 8量化实现331 GOPS的计算能力。系统吞吐量达到每秒196.7帧,输入图像大小为280 x 960。源代码发布在https://github.com/linbaiwpi/RoadNet-RT上。
In recent years, convolutional neural network (CNN) has gained popularity in many engineering applications especially for computer vision. In order to achieve better performance, more complex structures and advanced operations are incorporated into neural networks, which results in very long inference time. For time-critical tasks such as autonomous driving and virtual reality, real-time processing is fundamental. In order to reach real-time processing speed, a lightweight, high-throughput CNN architecture namely RoadNet-RT is proposed for road segmentation in this article. It achieves 92.55% MaxF score on KITTI road segmentation dataset. The inference time is about 9 ms per frame when running on GTX 1080 GPU. Comparing to the state-of-the-art network, RoadNet-RT speeds up the inference time by a factor of 17.8 at the cost of only 3.75% loss in accuracy. What is more, on CamVid dataset its accuracy is 92.98%. Several techniques such as depthwise separable convolution and non-uniformed kernel size convolution are optimized in the hardware accelerator design. The proposed CNN architecture has been successfully implemented on a ZCU102 MPSoC FPGA that achieves the computation capability of 331 GOPS using INT8 quantization. The system throughput reaches 196.7 frames per second with input image size of 280 x 960 . The source code is published at https://github.com/linbaiwpi/RoadNet-RT.