Dynamic Graph Warping Transformer for Video Alignment

Dynamic Graph Warping Transformer for Video Alignment
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Junyan Wang;Yang Long;M. Pagnucco;Yang Song
Junyan Wang;Yang Long;M. Pagnucco;Yang Song
中科院分区:
其他
文献类型:
--
作者:
Junyan Wang;Yang Long;M. Pagnucco;Yang Song

文献摘要

相似文献

视频对齐旨在匹配多个视频序列之间的同步动作信息。现有的方法通常基于监督学习来根据注释的动作阶段对齐视频帧。然而,这样的阶段级注释不能有效地引导帧级对齐,因为每个阶段可以跨个体以不同的速度完成。在本文中,我们引入动态扭曲考虑到视频之间的信息,一个新的动态图形扭曲变换器(DGWT)的网络模型。我们的方法是第一个专为视频分析和对齐而设计的图形Transformer框架。特别是,一种新的动态扭曲损失函数的设计,以对齐任意长度的视频使用注意力水平的功能。为了使邻接矩阵能够科普视频数据中的时间信息,提出了一种时间段图(TSG)。我们在两个公共数据集(Penn Action和Pouring)上的实验结果表明,与最先进的方法相比,该方法有了显著的改进。
Video alignment aims to match synchronised action information between multiple Video alignment aims to match synchronised action information between multiple video sequences. Existing methods are typically based on supervised learning to align video frames according to annotated action phases. However, such phase-level annotation cannot effectively guide frame-level alignment, since each phase can be completed at different speeds across individuals. In this paper, we introduce dynamic warping to take between-video information into account with a new Dynamic Graph Warping Trans-former (DGWT) network model. Our approach is the first Graph Transformer framework designed for video analysis and alignment. In particular, a novel dynamic warping loss function is designed to align videos of arbitrary length using attention-level features. A Temporal Segment Graph (TSG) is proposed to enable the adjacency matrix to cope with temporal information in video data. Our experimental results on two public datasets (Penn Action and Pouring) demonstrate significant improvements over state-of-the-art approaches.