CAREER: A Compression-Based Approach to Learning Video Representations
CAREER: A Compression-Based Approach to Learning Video Representations
批准号:
1845485
负责人:
Philipp Kraehenbuehl
金额:
$49.75万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
未结题
起止时间:
2019-06-01 至 2025-05-31
中文摘要
越来越多的数字通信、媒体消费和内容创作都围绕着视频展开。我们通过它们分享、观察和存档我们生活的许多方面。然而,设计和学习表征来理解这些视频被证明是具有挑战性的。直接将基于序列或图像的卷积神经网络扩展到视频只取得了适度的成功。该项目的目标是开发高效、健壮和紧凑的视频表示。该项目的压缩率每提高一个百分点,就意味着互联网流量的减少和存储效率的提高,从而降低了现代数字基础设施的巨大经济和环境成本。识别精度的任何提高都会带来更安全的自主代理,更灵敏的老年人监控和辅助技术,以及对体育和娱乐视频动态的更深入理解。此外,这项研究将通过更新和新的本科和研究生水平的视频识别和压缩课程转化为课堂。这个项目的技术目标分为四个重点。第一个重点是开发受视频压缩启发的视频识别模型。视频压缩社区开发了复杂、紧凑和高效的视频表示,用于存储大量数字媒体。该项目将研究视频压缩可以教给我们关于视频表示的东西,以及现代编解码器设计如何驱动深度视频模型的结构。第二个推力将视频识别的概念带回到压缩。压缩和识别之间的相互作用不是单行道。该项目将研究如何直接从数据中学习视频压缩,避开许多手动设计选择,以及如何学习视频压缩对丢失或损坏的信息具有鲁棒性。研究小组将开发一种新的视频压缩解释为重复的图像插值。这种解释为学习深度视频压缩算法打开了大门。第三个重点是研究用于识别和压缩任务的运动的光学表示。视频压缩和识别的核心都在于良好的运动表现。运动场将以紧凑、可压缩、时间一致和易于理解的方式表示。最后,第四个推力找到新的监督信号、评估任务及其相关数据。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
An ever-increasing amount of our digital communication, media consumption, and content creation revolves around videos. We share, watch, and archive many aspects of our lives through them. However, designing and learning representations to understand these videos has proven challenging. Direct extensions of sequence or image-based convolutional neural networks to videos have yielded only moderate success. The goal of this project is to develop efficient, robust, and compact video representations. Every percent increase in the compression rate from this project translates into decreased internet traffic and more storage efficiency, reducing the massive economic and environmental costs of modern digital infrastructure. Any increase in recognition accuracy results in safer autonomous agents, more responsive surveillance and assistive technologies for the elderly, and a deeper understanding of video dynamics in sports and entertainment. Furthermore, this research will translate to the classroom through updated and new undergraduate and graduate-level courses on video recognition and compression.The technical aim of this project is divided into four thrusts. The first thrust develops video recognition models inspired by video compression. The video compression community developed sophisticated, compact and efficient representations for video, used to store the bulk of digital media. The project will study what video compression can teach us about video representations, and how modern codec design can drive the structure of deep video models. The second thrust brings concepts from video recognition back to compression. The interplay between compression and recognition is not a one-way street. The project will investigate how video compression can be learned directly from data, side-stepping many of the manual design choices, and how video compression can learn to be robust to missing or corrupted information. The research team will develop a novel interpretation of video compression as repeated image interpolation. This interpretation opens the door to learned deep video compression algorithms. The third thrust studies the optical representation of motion for both recognition and compression tasks. At the core of both video compression and recognition lies a good representation of motion. The motion fields will be represented in a compact, compressible, temporally consistent, and easy to understand manner. Finally, the fourth thrust finds new supervisory signals, evaluation tasks, and their associated data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1007/978-3-031-20074-8_40
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
作者:
[Jang Hyun Cho;Philipp Krähenbühl]
通讯作者:
Jang Hyun Cho;Philipp Krähenbühl
DOI:
--
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
作者:
[Tianwei Yin;Xingyi Zhou;Philipp Krähenbühl]
通讯作者:
Tianwei Yin;Xingyi Zhou;Philipp Krähenbühl
DOI:
10.1109/iccv48922.2021.01530
发表时间:
2021-05
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
作者:
[Di Chen;V. Koltun;Philipp Krähenbühl]
通讯作者:
Di Chen;V. Koltun;Philipp Krähenbühl
DOI:
10.1109/cvpr42600.2020.00023
发表时间:
2019-12
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Chaoxia Wu;Ross B. Girshick;Kaiming He;Christoph Feichtenhofer;Philipp Krahenbuhl]
通讯作者:
Chaoxia Wu;Ross B. Girshick;Kaiming He;Christoph Feichtenhofer;Philipp Krahenbuhl
DOI:
10.1109/cvpr52729.2023.00691
发表时间:
2023-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
作者:
[Jang Hyun Cho]
通讯作者:
Jang Hyun Cho
共 10 条
RI: SMALL: Recognizing objects in images and their properties over time
-
批准号:2006820
-
项目类别:Standard Grant
-
资助金额:$41.95万
-
财政年份:2020
-
负责人:Philipp Kraehenbuehl
-
依托单位:
海外基金