MS-TCN plus plus : Multi-Stage Temporal Convolutional Network for Action Segmentation

MS-TCN plus plus : Multi-Stage Temporal Convolutional Network for Action Segmentation
复制标题

DOI:
10.1109/tpami.2020.3021756
复制
发表时间:
2023-06-01
影响因子:
23.6
通讯作者:
Gall, Juergen
Gall, Juergen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Shijie;Abu Farha, Yazan;Gall, Juergen

文献摘要

被引文献

相似文献

随着深度学习在对短修剪视频进行分类方面的成功,更多的注意力集中在对长未修剪视频中的活动进行时间分割和分类。最先进的动作分割方法利用几层时间卷积和时间池化。尽管这些方法在捕获时间依赖性的能力,他们的预测遭受过分割错误。在本文中,我们提出了一个多阶段的架构的时间动作分割任务,克服了以前的方法的局限性。第一阶段生成初始预测,该预测由下一阶段细化。在每个阶段中,我们堆叠几层膨胀的时间卷积,覆盖一个大的感受野,参数很少。虽然这种架构已经表现良好,但较低的层仍然受到小的接收场的影响。为了解决这个问题,我们提出了一个双扩张层,结合了大和小的感受野。我们进一步将第一阶段的设计与精炼阶段分离,以满足这些阶段的不同要求。广泛的评估表明,该模型在捕获长距离依赖关系和识别动作段的有效性。我们的模型在三个数据集上实现了最先进的结果:50沙拉,格鲁吉亚技术自我中心活动(GTEA)和早餐数据集。
With the success of deep learning in classifying short trimmed videos, more attention has been focused on temporally segmenting and classifying activities in long untrimmed videos. State-of-the-art approaches for action segmentation utilize several layers of temporal convolution and temporal pooling. Despite the capabilities of these approaches in capturing temporal dependencies, their predictions suffer from over-segmentation errors. In this paper, we propose a multi-stage architecture for the temporal action segmentation task that overcomes the limitations of the previous approaches. The first stage generates an initial prediction that is refined by the next ones. In each stage we stack several layers of dilated temporal convolutions covering a large receptive field with few parameters. While this architecture already performs well, lower layers still suffer from a small receptive field. To address this limitation, we propose a dual dilated layer that combines both large and small receptive fields. We further decouple the design of the first stage from the refining stages to address the different requirements of these stages. Extensive evaluation shows the effectiveness of the proposed model in capturing long-range dependencies and recognizing action segments. Our models achieve state-of-the-art results on three datasets: 50Salads, Georgia Tech Egocentric Activities (GTEA), and the Breakfast dataset.