Convolutional gated recurrent networks for video segmentation

Convolutional gated recurrent networks for video segmentation
复制标题

DOI:
10.1109/icip.2017.8296851
复制
发表时间:
2016-11
期刊:
2017 IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Mennatullah Siam;Sepehr Valipour;Martin Jägersand;Nilanjan Ray
Mennatullah Siam;Sepehr Valipour;Martin Jägersand;Nilanjan Ray
中科院分区:
其他
文献类型:
--
作者:
Mennatullah Siam;Sepehr Valipour;Martin Jägersand;Nilanjan Ray

文献摘要

被引文献

相似文献

语义分割最近取得了重大进展,但大多数以前的工作集中在改善单图像分割。在本文中,我们介绍了一种新的方法,隐式地利用视频中的时间数据进行在线分割。该设计接收一系列连续的视频帧,并输出最后一帧的分割。卷积门控递归网络用于递归部分,以保持图像中的空间连通性。该架构进行了测试的二进制和语义视频分割任务。实验在SegTrack V2、Davis、Camvid和Synthia的最新基准上进行。在我们所有的实验中,使用递归的全卷积网络提高了基线网络的性能。即,SegTrack 2和Davis的F-测量值分别改善5%和3%,Synthia和Camvid的平均IoU改善5.7%和1.6%。因此,RFCN网络可以被视为一种通过将其嵌入到利用时间数据的循环模块中来改进任何基线分割网络的方法。
Semantic segmentation has recently witnessed major progress, but most of the previous work focused on improving single image segmentation. In this paper, we introduce a novel approach to implicitly utilize temporal data in videos for online segmentation. This design receives a sequence of consecutive video frames and outputs the segmentation of the last frame. Convolutional gated recurrent networks are used for the recurrent part to preserve spatial connectivities in the image. This architecture is tested for both binary and semantic video segmentation tasks. Experiments are conducted on the recent benchmarks in SegTrack V2, Davis, Camvid, and Synthia. Using recurrent fully convolutional networks improved the baseline network performance in all of our experiments. Namely, 5% and 3% improvement of F-measure in SegTrack2 and Davis respectively, 5.7% and 1.6% improvement in mean IoU in Synthia and Camvid. Thus, RFCN networks can be seen as a method to improve any baseline segmentation network by embedding them into a recurrent module that utilizes temporal data.