Improving Action Segmentation via Graph-Based Temporal Reasoning

Improving Action Segmentation via Graph-Based Temporal Reasoning
复制标题

DOI:
10.1109/cvpr42600.2020.01404
复制
发表时间:
2020-06
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yifei Huang;Yusuke Sugano;Yoichi Sato
Yifei Huang;Yusuke Sugano;Yoichi Sato
中科院分区:
其他
文献类型:
--
作者:
Yifei Huang;Yusuke Sugano;Yoichi Sato

文献摘要

被引文献

相似文献

多个动作片段之间的时间关系在动作分割中起着重要作用,特别是当观察有限时(例如,动作被其他对象遮挡或发生在视场之外)。在本文中,我们提出了一个网络模块,称为基于图形的时间推理模块(GTRM),可以建立在现有的动作分割模型的顶部,学习在不同的时间跨度的多个动作段的关系。我们通过使用两个图卷积网络(GCN)对关系进行建模,其中每个节点代表一个动作段。这两个图具有不同的边缘属性,分别用于边界回归和分类任务。通过应用图卷积,我们可以根据每个节点与相邻节点的关系更新每个节点的表示。更新后的表示,然后用于改进的动作分割。我们评估我们的模型具有挑战性的自我中心的数据集,即EGTEA和EPIC-Kitterfly,其中的行动可能会部分观察到由于观点的限制。结果表明,我们提出的GTRM优于国家的最先进的动作分割模型的大幅度提高。我们还证明了我们的模型在两个第三人称视频数据集上的有效性,50 Salads数据集和Breakfast数据集。
Temporal relations among multiple action segments play an important role in action segmentation especially when observations are limited (e.g., actions are occluded by other objects or happen outside a field of view). In this paper, we propose a network module called Graph-based Temporal Reasoning Module (GTRM) that can be built on top of existing action segmentation models to learn the relation of multiple action segments in various time spans. We model the relations by using two Graph Convolution Networks (GCNs) where each node represents an action segment. The two graphs have different edge properties to account for boundary regression and classification tasks, respectively. By applying graph convolution, we can update each node's representation based on its relation with neighboring nodes. The updated representation is then used for improved action segmentation. We evaluate our model on the challenging egocentric datasets namely EGTEA and EPIC-Kitchens, where actions may be partially observed due to the viewpoint restriction. The results show that our proposed GTRM outperforms state-of-the-art action segmentation models by a large margin. We also demonstrate the effectiveness of our model on two third-person video datasets, the 50Salads dataset and the Breakfast dataset.