Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100

Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
复制标题

DOI:
10.1007/s11263-021-01531-2
复制
发表时间:
2021-10-20
影响因子:
19.5
通讯作者:
Wray, Michael
Wray, Michael
中科院分区:
计算机科学2区
文献类型:
--
作者:
Damen, Dima;Doughty, Hazel;Wray, Michael

文献摘要

被引文献

相似文献

本文介绍了扩展自我中心视觉中最大数据集EPIC-KITCHENS的管道。EPIC-KITCHENS-100是一个100小时、2000万帧、90000个动作的700个可变长度视频的集合,使用头戴式摄像机在45个环境中捕捉长期的无脚本活动。与之前的版本(Damen in Scaling egocentric vision:ECCV,2018)相比,EPIC-KITCHENS-100使用了一种新的管道进行了注释,该管道允许更密集(每分钟增加54%的动作)和更完整的细粒度动作注释(增加128%的动作片段)。这些数据集带来了新的挑战,例如动作检测和评估“时间测试”,即在2018年收集的数据上训练的模型是否可以推广到两年后收集的新镜头。该数据集与6个挑战保持一致:动作识别(完全和弱监督),动作检测,动作预测,跨模式检索(从字幕),以及用于动作识别的无监督域自适应。对于每个挑战,我们定义任务,提供基线和评估指标。
This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing long-term unscripted activities in 45 environments, using head-mounted cameras. Compared to its previous version (Damen in Scaling egocentric vision: ECCV, 2018), EPIC-KITCHENS-100 has been annotated using a novel pipeline that allows denser (54% more actions per minute) and more complete annotations of fine-grained actions (+128% more action segments). This collection enables new challenges such as action detection and evaluating the "test of time"-i.e. whether models trained on data collected in 2018 can generalise to new footage collected two years later. The dataset is aligned with 6 challenges: action recognition (full and weak supervision), action detection, action anticipation, cross-modal retrieval (from captions), as well as unsupervised domain adaptation for action recognition. For each challenge, we define the task, provide baselines and evaluation metrics.