Sparse Representations for Object- and Ego-Motion Estimations in Dynamic Scenes

Sparse Representations for Object- and Ego-Motion Estimations in Dynamic Scenes
复制标题

DOI:
10.1109/tnnls.2020.3006467
复制
发表时间:
2020-07
影响因子:
10.4
通讯作者:
H. Kashyap;Charless C. Fowlkes;J. Krichmar
H. Kashyap;Charless C. Fowlkes;J. Krichmar
中科院分区:
计算机科学1区
文献类型:
--
作者:
H. Kashyap;Charless C. Fowlkes;J. Krichmar

文献摘要

相似文献

在自主运动或自我运动过程中,解开动态场景中视觉运动的来源对于自主导航和跟踪非常重要。在包含独立运动物体的视频帧的动态图像片段中,相对于下一帧的光流是由于摄像机和物体运动而产生的运动场的总和。传统的自我运动估计方法假设场景是静态的,而最近的基于深度学习的方法没有将像素速度分离为对象和自我运动分量。我们提出了一种基于学习的方法,使用卷积自动编码器从图像序列中预测自我运动参数和对象运动场(OMF),同时对不受约束的场景深度引起的变化具有鲁棒性。这是通过以下方式实现的:1)利用允许独立于深度求解自我运动参数的连续自我运动约束进行训练,以及2)学习稀疏激活的过完备自我运动场(EMF)基集,其消除了用于自我运动估计任务的静态和动态段中的不相关分量。为了学习EMF基集,我们提出了一个新的可微稀疏罚函数,它近似于自动编码器瓶颈层中非零激活的数量,并且比基于L1和L2范数的罚函数更有效地执行稀疏性。与现有的直接自我运动估计方法不同,预测的全局EMF可以通过将其与光流进行比较来直接用于提取OMF。与最先进的基线相比,所提出的模型在动态场景的真实的和合成数据集上进行评估时,在逐像素对象和自我运动估计任务上表现良好。
Disentangling the sources of visual motion in a dynamic scene during self-movement or ego motion is important for autonomous navigation and tracking. In the dynamic image segments of a video frame containing independently moving objects, optic flow relative to the next frame is the sum of the motion fields generated due to camera and object motion. The traditional ego-motion estimation methods assume the scene to be static, and the recent deep learning-based methods do not separate pixel velocities into object- and ego-motion components. We propose a learning-based approach to predict both ego-motion parameters and object-motion field (OMF) from image sequences using a convolutional autoencoder while being robust to variations due to the unconstrained scene depth. This is achieved by: 1) training with continuous ego-motion constraints that allow solving for ego-motion parameters independently of depth and 2) learning a sparsely activated overcomplete ego-motion field (EMF) basis set, which eliminates the irrelevant components in both static and dynamic segments for the task of ego-motion estimation. In order to learn the EMF basis set, we propose a new differentiable sparsity penalty function that approximates the number of nonzero activations in the bottleneck layer of the autoencoder and enforces sparsity more effectively than L1- and L2-norm-based penalties. Unlike the existing direct ego-motion estimation methods, the predicted global EMF can be used to extract OMF directly by comparing it against the optic flow. Compared with the state-of-the-art baselines, the proposed model performs favorably on pixelwise object- and ego-motion estimation tasks when evaluated on real and synthetic data sets of dynamic scenes.