Deep learning of spatio-temporal features with geometric-based moving point detection for motion segmentation

Deep learning of spatio-temporal features with geometric-based moving point detection for motion segmentation
复制标题

DOI:
10.1109/icra.2014.6907299
复制
发表时间:
2014-05
期刊:
2014 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Tsung-Han Lin;C. Wang
Tsung-Han Lin;C. Wang
中科院分区:
其他
文献类型:
--
作者:
Tsung-Han Lin;C. Wang

文献摘要

被引文献

相似文献

本文介绍了一种基于深度学习的移动立体相机完成运动分割的方法。先前的运动目标检测工作大多使用基于 3D 几何约束的点特征。然而,点特征需要良好的特征,并且在物体具有平滑纹理的情况下难以检测或正确匹配。为了缓解这个问题,提出了基于重建独立分量分析(RICA)自动编码器从原始图像数据无监督地学习高级时空特征。尽管新的时空特征很强大,但这些特征不能不学习并用于解释动态场景的 3D 几何形状,这对于移动摄像机的移动物体检测至关重要。由于基于 3D 几何约束检测到的移动点仍然包含 3D 场景以及相机自我运动的有价值的信息,因此我们提出了一个框架,该框架将检测到的移动点结果和学习到的时空特征结合起来,作为执行运动分割的递归神经网络 (RNN) 的输入。这两个功能有效地相辅相成。所提出的方法通过包含多个移动对象的真实立体视频数据进行了演示,并且与现有的基于 3D 几何的移动点检测器相比,检测率提高了 26%。
This paper introduces an approach to accomplish motion segmentation from a moving stereo camera based on deep learning. Previous work on moving object detection mostly use point features based on 3D geometric constraints. However, point features require good features, and are hard to detect or to be matched correctly in situations where objects have smooth textures. To alleviate this problem, learning high-level spatio-temporal features unsupervisedly from raw image data based on Reconstruction Independent Component Analysis (RICA) autoencoders is proposed. Despite the power of the new spatio-temporal features, these features cannot not learn and be used to interpret 3D geometry of dynamic scenes, which is critical for moving object detection from moving cameras. As detected moving points based on 3D geometric constraints still contain valuable information of 3D scene as well as the camera egomotion, we propose a framework that incorporates both the detected moving point results and the learned spatio-temporal features as inputs to Recursive Neural Networks (RNN) that performs motion segmentation. Both features effectively complement each other. The proposed approach is demonstrated with real-world stereo video data that contains multiple moving objects, and has achieved 26% better detection rate over the existing 3D geometric-based moving points detector.