Noisy-LSTM: Improving Temporal Awareness for Video Semantic Segmentation

Noisy-LSTM: Improving Temporal Awareness for Video Semantic Segmentation
复制标题

DOI:
10.1109/access.2021.3067928
复制
发表时间:
2020-10
期刊:
影响因子:
3.9
通讯作者:
Bowen Wang;Liangzhi Li;Yuta Nakashima;R. Kawasaki;H. Nagahara;Y. Yagi
Bowen Wang;Liangzhi Li;Yuta Nakashima;R. Kawasaki;H. Nagahara;Y. Yagi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bowen Wang;Liangzhi Li;Yuta Nakashima;R. Kawasaki;H. Nagahara;Y. Yagi

文献摘要

相似文献

语义视频分割是各种应用的关键挑战。本文提出了一种名为noise - lstm的新模型,该模型可以端到端训练,使用卷积lstm (convlstm)来利用视频帧中的时间相干性,以及一种简单而有效的训练策略,即用噪声替换给定视频序列中的帧。我们的训练策略破坏了视频帧的时间相干性,从而使ConvLSTMs中的时间链接不可靠;因此,这可能会提高模型从视频帧中提取特征的能力,并作为正则化器避免过拟合,而不需要额外的数据注释或计算成本。实验结果表明,该模型可以在cityscape和EndoVis2018数据集上实现最先进的性能。所建议的方法的代码可在https://github.com/wbw520/NoisyLSTM上获得。
Semantic video segmentation is a key challenge for various applications. This paper presents a new model named Noisy-LSTM, which is trainable in an end-to-end manner, with convolutional LSTMs (ConvLSTMs) to leverage the temporal coherence in video frames, together with a simple yet effective training strategy that replaces a frame in a given video sequence with noises. Our training strategy spoils the temporal coherence in video frames and thus makes the temporal links in ConvLSTMs unreliable; this may consequently improve the ability of the model to extract features from video frames and serve as a regularizer to avoid overfitting, without requiring extra data annotations or computational costs. Experimental results demonstrate that the proposed model can achieve state-of-the-art performances on both the CityScapes and EndoVis2018 datasets. The code for the proposed method is available at https://github.com/wbw520/NoisyLSTM.