CAREER: Representing, Discovering, and Assembling Motifs for Video Understanding
CAREER: Representing, Discovering, and Assembling Motifs for Video Understanding
批准号:
2238769
负责人:
Abhinav Shrivastava
金额:
$59.81万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-15 至 2028-01-31
中文摘要
该项目将建立创新的技术,使计算机能够理解时间现象,如视频中的人类行为,并有可能改变安全、卫生和机器人领域的应用程序。时间现象在不同的时间尺度上表现出结构--长事件(例如早餐)由多个长期活动(例如烹饪煎蛋卷)组成,而这些长期活动又由各种原子动作(例如切洋葱)组成。这个项目将开发时间现象的计算表示法,通过强调视频中的时间线索来捕捉概念如何随着时间的推移而演变。利用这些表示法,这项工作将创建软件,可以从视频集合中发现独特的反复出现的原子动作,并学习组成这些原子动作,以理解和解释长期、复杂的时间现象。该项目的成果将使计算机能够更好地识别复杂的活动、检测异常和预测未来的行动。开发的技术有可能解决几个领域的挑战,例如分析天气模式和预测极端事件,在互联网上搜索恶意内容的来源,以及使用自然活动级别的查询对互联网规模的视频集合进行索引和搜索。与这项研究相结合的是一项全面的教育、指导和推广计划,包括在多个层面上培训学生进行研究,为本科生和研究生课程的发展做出贡献,并设计推广计划以吸引多个层面的不同学生。在技术层面上,该项目将解决理解长期时间现象的根本挑战。这个项目的核心是“主题”的概念--独特的重复的时间模式--可以组合成长期的叙述,比如活动。该研究计划寻求在三个关键领域取得进展:(A)无监督时间表征学习,这项工作将开发能够更好地对时间现象进行建模的分离的时间表征;(B)发现主题,这将开发一个大规模框架,用于从未标记的视频中发现独特的重复时间模式,作为原子动作,这可以帮助理解长期活动;以及(C)学习组装主题来分解动作。该项目提出了可扩展的策略,以从未标记的视频中学习隐式和显式的随机动作语法,使用当代的数据驱动方法重新探讨了一个长期存在的问题。这项研究工作为视频理解领域令人兴奋的新研究方向提供了路线图。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project will build innovative technology to allow computers to understand temporal phenomena, such as human actions in videos, with the potential of transforming applications across security, health, and robotics. Temporal phenomena exhibit structure at various time scales—long events (e.g., breakfast) are composed of multiple long-term activities (e.g., cooking an omelet), which in turn are composed of various atomic actions (e.g., cutting onions). This project will develop computational representations of temporal phenomena that capture how concepts evolve over time by accentuating temporal cues in videos. Leveraging these representations, this work will create software that can discover distinctive recurring atomic actions from video collections and learn to compose these atomic actions to understand and explain long-term, complex temporal phenomena. The outcomes of this project will enable computers to be significantly better at recognizing complex activities, detecting anomalies, and forecasting future actions. The developed technologies have the potential to address challenges in several areas, such as analyzing weather patterns and forecasting extreme events, provenance search for malicious content on the internet, and indexing and searching internet-scale video collections using natural activity-level queries. Integrated with the research is a comprehensive plan for education, mentoring, and outreach, including training students in research at multiple levels, contributing to curriculum development for undergraduate and graduate courses, and designing outreach programs to attract diverse students at multiple levels.At a technical level, this project will address fundamental challenges in understanding long-term temporal phenomena. At the core of this project is the notion of ‘motifs’—distinctive repeating temporal patterns—that can be assembled into long-term narratives, such as activities. The research program seeks advances in three key areas: (a) unsupervised temporal representation learning, where this work will develop disentangled temporal representations that can better model temporal phenomena, (b) discovering motifs, which will develop a large scale framework for discovering distinctive repeating temporal patterns as atomic actions, from unlabeled videos, which can help understand long-term activities, and (c) learning to assemble motifs to decompose actions. The project presents scalable strategies to learn implicit and explicit stochastic grammars of actions from unlabeled videos, revisiting a long-standing problem using contemporary data-driven methods. This research effort provides a roadmap for exciting new research directions in video understanding.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/iccv51070.2023.01852
发表时间:
2023-09
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
作者:
[Nirat Saini;Hanyu Wang;Archana Swaminathan;Vinoj Jayasundara;Bo He;Kamal Gupta;Abhinav Shrivastava]
通讯作者:
Nirat Saini;Hanyu Wang;Archana Swaminathan;Vinoj Jayasundara;Bo He;Kamal Gupta;Abhinav Shrivastava
DOI:
10.1109/iccv51070.2023.00382
发表时间:
2023-03
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
作者:
[Kamal Gupta;V. Jampani;Carlos Esteves;Abhinav Shrivastava;A. Makadia;Noah Snavely;Abhishek Kar]
通讯作者:
Kamal Gupta;V. Jampani;Carlos Esteves;Abhinav Shrivastava;A. Makadia;Noah Snavely;Abhishek Kar
海外基金