Discriminative Feature Learning for Video Semantic Segmentation

Discriminative Feature Learning for Video Semantic Segmentation
复制标题

DOI:
10.1109/icvrv.2014.65
复制
发表时间:
2014-08
期刊:
2014 International Conference on Virtual Reality and Visualization
影响因子:
--
通讯作者:
Han Zhang;Kai Jiang;Yu Zhang;Qing Li;Changqun Xia;Xiaowu Chen
Han Zhang;Kai Jiang;Yu Zhang;Qing Li;Changqun Xia;Xiaowu Chen
中科院分区:
其他
文献类型:
--
作者:
Han Zhang;Kai Jiang;Yu Zhang;Qing Li;Changqun Xia;Xiaowu Chen

文献摘要

被引文献

相似文献

本文提出了一种新的基于深度学习的视频语义分割方法。特别地,我们利用三维卷积神经网络(3D CNN)从时空体中学习判别层次特征,以实现准确的像素标记。学习到的特征能够捕获外观和运动信息。为了沿着真实对象边界对齐像素标签,并保持局部一致性,我们进一步对从输入视频中提取的相干3d区域或超级体素构建的图执行图切。实验表明,由于学习特征的判别能力,即使训练样本很少,我们的方法也可以在没有复杂推理模型的情况下获得与现有方法相比具有竞争力的标签准确性。
In this paper, we propose a novel deep learning based method for video semantic segmentation. Specially, we utilize 3D convolution neural network (3D CNN) to learn discriminative hierarchical features from spatial-temporal volumes for accurate pixel labelling. The learned features are capable of capturing both appearance and motion information. To align the pixel labels along real object boundaries, as well as maintain local consistency, we further perform graph-cut on a graph constructed on coherent 3d regions, or super-voxels, extracted from input video. Experiments demonstrate that due to the discriminative capability of learned features, our approach can obtain competitive labelling accuracy compared to the state-of-art in absence of sophisticated inference models, even with few training samples.