Learning Temporal Sentence Grounding From Narrated EgoVideos

Learning Temporal Sentence Grounding From Narrated EgoVideos
复制标题

DOI:
10.48550/arxiv.2310.17395
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Kevin Flanagan;D. Damen;Michael Wray
Kevin Flanagan;D. Damen;Michael Wray
中科院分区:
其他
文献类型:
--
作者:
Kevin Flanagan;D. Damen;Michael Wray

文献摘要

相似文献

长格式自我中心数据集(如Ego 4D和EPIC-Kitchenette)的出现对时间句子基础(TSG)的任务提出了新的挑战。与评估此任务的传统基准相比,这些数据集提供了更细粒度的句子,以支持更长的视频。在本文中,我们开发了一种方法,用于学习在这些数据集中只使用叙述及其相应的粗略叙述时间戳的基础句子。我们建议人工合并剪辑,以对比的方式使用文本调节注意训练时间接地。与高性能的TSG方法相比,这种剪辑合并(CliMer)方法被证明是有效的-例如,Ego 4D上的平均R@1从3.9提高到5.7,EPIC-Kitchild上的平均R@1从10.7提高到13.0。代码和数据分割可从https://github.com/keflanagan/CliMer获得
The onset of long-form egocentric datasets such as Ego4D and EPIC-Kitchens presents a new challenge for the task of Temporal Sentence Grounding (TSG). Compared to traditional benchmarks on which this task is evaluated, these datasets offer finer-grained sentences to ground in notably longer videos. In this paper, we develop an approach for learning to ground sentences in these datasets using only narrations and their corresponding rough narration timestamps. We propose to artificially merge clips to train for temporal grounding in a contrastive manner using text-conditioning attention. This Clip Merging (CliMer) approach is shown to be effective when compared with a high performing TSG method -- e.g. mean R@1 improves from 3.9 to 5.7 on Ego4D and from 10.7 to 13.0 on EPIC-Kitchens. Code and data splits available from: https://github.com/keflanagan/CliMer