VehiPose: a multi-scale framework for vehicle pose estimation

VehiPose: a multi-scale framework for vehicle pose estimation
复制标题

DOI:
10.1117/12.2595800
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Divyanshu Gupta;Bruno Artacho;A. Savakis
Divyanshu Gupta;Bruno Artacho;A. Savakis
中科院分区:
其他
文献类型:
--
作者:
Divyanshu Gupta;Bruno Artacho;A. Savakis

文献摘要

相似文献

车辆姿态估计对于诸如自动驾驶汽车、交通监控和场景分析之类的应用是有用的。计算机视觉和深度学习的最新发展在人体姿态估计方面取得了重大进展,但这项工作很少应用于车辆姿态。我们提出了VehiPose,这是一种有效的车辆姿态估计架构,基于多尺度深度学习方法,可实现高精度车辆姿态估计,同时保持可管理的网络复杂性和模块化。VehiPose架构将编码器-解码器架构与瀑布式卷积模块相结合,用于多尺度特征表示。我们的方法旨在减少由于连续池层的损失,并保留编码器特征表示中的多尺度上下文和空间信息。瀑布模块生成多尺度特征,因为它利用渐进式过滤的效率,同时通过多个特征的串联来保持更宽的视野。这种多尺度方法产生了一种鲁棒的车辆姿态估计架构,该架构包含跨尺度的上下文信息,并在端到端可训练网络中执行车辆关键点的定位。
Vehicle pose estimation is useful for applications such as self-driving cars, traffic monitoring, and scene analysis. Recent developments in computer vision and deep learning have achieved significant progress in human pose estimation, but little of this work has been applied to vehicle pose. We propose VehiPose, an efficient architecture for vehicle pose estimation, based on a multi-scale deep learning approach that achieves high accuracy vehicle pose estimation while maintaining manageable network complexity and modularity. The VehiPose architecture combines an encoder-decoder architecture with a waterfall atrous convolution module for multi-scale feature representation. Our approach aims to reduce the loss due to successive pooling layers and preserve the multiscale contextual and spatial information in the encoder feature representations. The waterfall module generates multiscale features, as it leverages the efficiency of progressive filtering while maintaining wider fields-of-view through the concatenation of multiple features. This multi-scale approach results in a robust vehicle pose estimation architecture that incorporates contextual information across scales and performs the localization of vehicle keypoints in an end-to-end trainable network.