Ego4D: Around the World in 3,000 Hours of Egocentric Video

Ego4D: Around the World in 3,000 Hours of Egocentric Video
复制标题

DOI:
10.1109/cvpr52688.2022.01842
复制
发表时间:
2021-10
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
K. Grauman;Andrew Westbury;Eugene Byrne;Zachary Chavis;Antonino Furnari;Rohit Girdhar;Jackson Hamburger;Hao Jiang;Miao Liu;Xingyu Liu;Miguel Martin;Tushar Nagarajan;Ilija Radosavovic;Santhosh K. Ramakrishnan;Fiona Ryan;J. Sharma;Michael Wray;Mengmeng Xu;Eric Z. Xu;Chen Zhao;Siddhant Bansal;Dhruv Batra;Vincent Cartillier;S. Crane;Tien Do;Morrie Doulaty;Akshay Erapalli;Christoph Feichtenhofer;A. Fragomeni;Qichen Fu;Christian Fuegen;A. Gebreselasie;Cristina González;James M. Hillis;Xuhua Huang;Yifei Huang;Wenqi Jia;Weslie Khoo;J. Kolár;Satwik Kottur;Anurag Kumar;F. Landini;Chao Li;Yanghao Li;Zhenqiang Li;K. Mangalam;Raghava Modhugu;Jonathan Munro;Tullie Murrell;Takumi Nishiyasu;Will Price;Paola Ruiz Puentes;Merey Ramazanova;Leda Sari;K. Somasundaram;Audrey Southerland;Yusuke Sugano;Ruijie Tao;Minh Vo;Yuchen Wang;Xindi Wu;Takuma Yagi;Yunyi Zhu;P. Arbeláez;David J. Crandall;D. Damen;G. Farinella;Bernard Ghanem;V. Ithapu;C. V. Jawahar;H. Joo;Kris Kitani;Haizhou Li;Richard A. Newcombe;A. Oliva;H. Park;James M. Rehg;Yoichi Sato;Jianbo Shi;Mike Zheng Shou;A. Torralba;L. Torresani;Mingfei Yan;J. Malik
K. Grauman;Andrew Westbury;Eugene Byrne;Zachary Chavis;Antonino Furnari;Rohit Girdhar;Jackson Hamburger;Hao Jiang;Miao Liu;Xingyu Liu;Miguel Martin;Tushar Nagarajan;Ilija Radosavovic;Santhosh K. Ramakrishnan;Fiona Ryan;J. Sharma;Michael Wray;Mengmeng Xu;Eric Z. Xu;Chen Zhao;Siddhant Bansal;Dhruv Batra;Vincent Cartillier;S. Crane;Tien Do;Morrie Doulaty;Akshay Erapalli;Christoph Feichtenhofer;A. Fragomeni;Qichen Fu;Christian Fuegen;A. Gebreselasie;Cristina González;James M. Hillis;Xuhua Huang;Yifei Huang;Wenqi Jia;Weslie Khoo;J. Kolár;Satwik Kottur;Anurag Kumar;F. Landini;Chao Li;Yanghao Li;Zhenqiang Li;K. Mangalam;Raghava Modhugu;Jonathan Munro;Tullie Murrell;Takumi Nishiyasu;Will Price;Paola Ruiz Puentes;Merey Ramazanova;Leda Sari;K. Somasundaram;Audrey Southerland;Yusuke Sugano;Ruijie Tao;Minh Vo;Yuchen Wang;Xindi Wu;Takuma Yagi;Yunyi Zhu;P. Arbeláez;David J. Crandall;D. Damen;G. Farinella;Bernard Ghanem;V. Ithapu;C. V. Jawahar;H. Joo;Kris Kitani;Haizhou Li;Richard A. Newcombe;A. Oliva;H. Park;James M. Rehg;Yoichi Sato;Jianbo Shi;Mike Zheng Shou;A. Torralba;L. Torresani;Mingfei Yan;J. Malik
中科院分区:
其他
文献类型:
--
作者:
K. Grauman;Andrew Westbury;Eugene Byrne;Zachary Chavis;Antonino Furnari;Rohit Girdhar;Jackson Hamburger;Hao Jiang;Miao Liu;Xingyu Liu;Miguel Martin;Tushar Nagarajan;Ilija Radosavovic;Santhosh K. Ramakrishnan;Fiona Ryan;J. Sharma;Michael Wray;Mengmeng Xu;Eric Z. Xu;Chen Zhao;Siddhant Bansal;Dhruv Batra;Vincent Cartillier;S. Crane;Tien Do;Morrie Doulaty;Akshay Erapalli;Christoph Feichtenhofer;A. Fragomeni;Qichen Fu;Christian Fuegen;A. Gebreselasie;Cristina González;James M. Hillis;Xuhua Huang;Yifei Huang;Wenqi Jia;Weslie Khoo;J. Kolár;Satwik Kottur;Anurag Kumar;F. Landini;Chao Li;Yanghao Li;Zhenqiang Li;K. Mangalam;Raghava Modhugu;Jonathan Munro;Tullie Murrell;Takumi Nishiyasu;Will Price;Paola Ruiz Puentes;Merey Ramazanova;Leda Sari;K. Somasundaram;Audrey Southerland;Yusuke Sugano;Ruijie Tao;Minh Vo;Yuchen Wang;Xindi Wu;Takuma Yagi;Yunyi Zhu;P. Arbeláez;David J. Crandall;D. Damen;G. Farinella;Bernard Ghanem;V. Ithapu;C. V. Jawahar;H. Joo;Kris Kitani;Haizhou Li;Richard A. Newcombe;A. Oliva;H. Park;James M. Rehg;Yoichi Sato;Jianbo Shi;Mike Zheng Shou;A. Torralba;L. Torresani;Mingfei Yan;J. Malik

文献摘要

被引文献

相似文献

我们介绍了EGO4D,这是一个大规模的自我中心视频数据集和基准套件。它提供了3,670小时的Dailylife活动视频,这些视频涵盖了数百个场景(家庭,户外,工作场所,休闲等),这些场景由来自74个世界各地和9个不同国家 /地区的931个独特的摄像头佩戴者捕获。收集方法旨在维护严格的隐私和道德标准,并在相关的情况下同意参与者和强大的去识别程序。 EGO4D大大扩展了研究界公开可用的各种以自我为中心的视频素材的数量。视频的某些部分伴随着音频,环境的3D网格,眼睛凝视,立体声和/或在同一事件中来自多个以egentric摄像机的同步视频。此外,我们提出了许多新的基准测试挑战,这些挑战围绕了解过去的第一人称视觉体验(查询情节记忆),现在(分析手动操纵,视听对话和社交互动)以及未来(预测活动)。通过公开共享这个大规模的注释数据集和基准套件,我们旨在推动第一人称感知的边界。项目页面:https://ego4d-data.org/
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/