STFCN: Spatio-Temporal Fully Convolutional Neural Network for Semantic Segmentation of Street Scenes

STFCN: Spatio-Temporal Fully Convolutional Neural Network for Semantic Segmentation of Street Scenes
复制标题

DOI:
10.1007/978-3-319-54407-6_33
复制
发表时间:
2016-11
期刊:
--
影响因子:
--
通讯作者:
Mohsen Fayyaz-;M. H. Saffar;M. Sabokrou;M. Fathy;F. Huang;R. Klette
Mohsen Fayyaz-;M. H. Saffar;M. Sabokrou;M. Fathy;F. Huang;R. Klette
中科院分区:
其他
文献类型:
--
作者:
Mohsen Fayyaz-;M. H. Saffar;M. Sabokrou;M. Fathy;F. Huang;R. Klette

文献摘要

被引文献

相似文献

本文提出了一种新的方法,涉及的空间和时间特征的语义分割的街道场景。目前对卷积神经网络(CNN)的研究表明,CNN提供了先进的空间特征,为语义分割任务提供了非常好的解决方案。我们调查如何涉及时间特征也有很好的效果分割视频数据。我们提出了一个基于循环神经网络的沿着短期记忆(LSTM)架构的模块,用于解释视频帧随时间的时间特性。我们的系统将视频的帧作为输入,并产生相应大小的输出;为了分割视频,我们的方法结合了三个组件的使用:首先,使用CNN提取帧的区域空间特征;然后,使用LSTM添加时间特征;最后,通过对时空特征进行去卷积,我们产生像素预测。我们的关键见解是构建时空卷积网络(时空CNN),它具有用于语义视频分割的端到端架构。我们完全适应了一些已知的卷积网络架构(如FCN-AlexNet和FCN-VGG 16),并将卷积扩展到我们的时空CNN中。我们的时空CNN实现了最先进的语义分割,如Camvid和NYUDv 2数据集所示。
This paper presents a novel method to involve both spatial and temporal features for semantic segmentation of street scenes. Current work onconvolutional neural networks(CNNs) has shown that CNNs provide advanced spatial features supporting a very good performance of solutions for the semantic segmentation task. We investigate how involving temporal features also has a good effect on segmenting video data. We propose a module based on along short-term memory(LSTM) architecture of a recurrent neural network for interpreting the temporal characteristics of video frames over time. Our system takes as input frames of a video and produces a correspondingly-sized output; for segmenting the video our method combines the use of three components: First, the regional spatial features of frames are extracted using a CNN; then, using LSTM the temporal features are added; finally, by deconvolving the spatio-temporal features we produce pixel-wise predictions. Our key insight is to buildspatio-temporal convolutional networks(spatio-temporal CNNs) that have an end-to-end architecture for semantic video segmentation. We adapted fully some known convolutional network architectures (such as FCN-AlexNet and FCN-VGG16), and dilated convolution into our spatio-temporal CNNs. Our spatio-temporal CNNs achieve state-of-the-art semantic segmentation, as demonstrated for the Camvid and NYUDv2 datasets.