Algorithm and VLSI Architecture Co-Design on Efficient Semi-Global Stereo Matching

Algorithm and VLSI Architecture Co-Design on Efficient Semi-Global Stereo Matching
复制标题

DOI:
10.1109/tcsvt.2019.2957275
复制
发表时间:
2020-11
影响因子:
8.4
通讯作者:
Xuchong Zhang;He Dai;Hongbin Sun;Nanning Zheng
Xuchong Zhang;He Dai;Hongbin Sun;Nanning Zheng
中科院分区:
工程技术1区
文献类型:
--
作者:
Xuchong Zhang;He Dai;Hongbin Sun;Nanning Zheng

文献摘要

相似文献

半全局匹配(Semi-global matching, SGM)由于在视差图像质量和计算复杂度之间取得了很好的平衡,在高精度实时立体匹配设计中受到青睐。然而,到目前为止,大多数SGM设计都局限于小图像分辨率和视差范围的实时处理,或者以视差图像质量显著下降为代价,通过简化原始算法来实现高吞吐量。我们分析了高效SGM设计的主要挑战是其存储器架构,包括片上存储器成本和片外存储器带宽。我们通过算法和架构协同设计来解决内存架构的挑战。基于观察到的SGM算法的不完备性和不准确性两个特点,本文分别提出了降低片上存储器成本和压缩片外存储器带宽的几种有效技术。此外,我们还设计了高吞吐量和流水线架构来实现所提出的技术。在KITTI2015和Middlebury V3立体数据集上对SGM设计的视差图像质量和硬件效率进行了评估。评估结果表明,与最佳参考设计技术相比,所提出的电路设计在视差范围为128时的吞吐量可以轻松达到1080P@30fps,并且可以在获得更好或相同视差图像质量的同时,将片上存储器成本和片外存储器带宽分别降低高达4倍和2倍。
Semi-global matching (SGM) is favored for high accuracy real-time stereo matching design as it achieves a good trade-off between disparity image quality and computational complexity. Nevertheless, most of previous SGM designs so far are restricted to the real-time processing of small image resolution and disparity range, or achieve high throughput by simplifying the original algorithm at the penalty of significant disparity image quality degradation. We analyze that the major challenge to efficient SGM design is its memory architecture, including both on-chip memory cost and off-chip memory bandwidth. We address the memory architecture challenge by algorithm and architecture co-design. Based on two observed features of SGM algorithm, i.e. incompleteness and inaccuracy, this paper proposes several efficient techniques to reduce on-chip memory cost and compress off-chip memory bandwidth respectively. Moreover, we also design high throughput and pipelined architecture to implement the proposed techniques. The disparity image quality and hardware efficiency of the proposed SGM design are evaluated on both KITTI2015 and Middlebury V3 stereo datasets. Evaluation results demonstrate that, the throughput of the proposed circuit designs can easily achieve 1080P@30fps at the disparity range of 128, and can reduce the on-chip memory cost and off-chip memory bandwidth by up to $4\times $ and $2\times $ respectively while achieving better or the same disparity image quality, compared with the best reference design techniques.