Spotting Temporally Precise, Fine-Grained Events in Video

Spotting Temporally Precise, Fine-Grained Events in Video
复制标题

DOI:
10.48550/arxiv.2207.10213
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
James Hong;Haotian Zhang;Michaël Gharbi;Matthew Fisher;Kayvon Fatahalian
James Hong;Haotian Zhang;Michaël Gharbi;Matthew Fisher;Kayvon Fatahalian
中科院分区:
其他
文献类型:
--
作者:
James Hong;Haotian Zhang;Michaël Gharbi;Matthew Fisher;Kayvon Fatahalian

文献摘要

相似文献

我们介绍了在视频中发现时间精确、细粒度事件的任务(检测事件发生的精确时刻)。精确定位需要模型在全局范围内推断出动作的全时规模,并在局部范围内识别出细微的帧与帧之间的外观和动作差异,从而识别这些动作中的事件。令人惊讶的是,我们发现对先前的视频理解任务(如动作检测和分割)的最佳解决方案并不能同时满足这两个要求。作为回应,我们提出了E2E-Spot,这是一个紧凑的端到端模型,在精确定位任务上表现良好,可以在单个GPU上快速训练。我们证明E2E-Spot显著优于最近从视频动作检测、分割和定位文献中适应的基线,以精确定位任务。最后,我们为几个细粒度的体育动作数据集提供了新的注释和分割,以使这些数据集适合未来的精确定位工作。
We introduce the task of spotting temporally precise, fine-grained events in video (detecting the precise moment in time events occur). Precise spotting requires models to reason globally about the full-time scale of actions and locally to identify subtle frame-to-frame appearance and motion differences that identify events during these actions. Surprisingly, we find that top performing solutions to prior video understanding tasks such as action detection and segmentation do not simultaneously meet both requirements. In response, we propose E2E-Spot, a compact, end-to-end model that performs well on the precise spotting task and can be trained quickly on a single GPU. We demonstrate that E2E-Spot significantly outperforms recent baselines adapted from the video action detection, segmentation, and spotting literature to the precise spotting task. Finally, we contribute new annotations and splits to several fine-grained sports action datasets to make these datasets suitable for future work on precise spotting.