MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Zelun Luo;Zane Durante;Linden Li;Wanze Xie;Ruochen Liu;Emily Jin;Zhuoyi Huang;Lun Yu Li
Zelun Luo;Zane Durante;Linden Li;Wanze Xie;Ruochen Liu;Emily Jin;Zhuoyi Huang;Lun Yu Li
中科院分区:
其他
文献类型:
--
作者:
Zelun Luo;Zane Durante;Linden Li;Wanze Xie;Ruochen Liu;Emily Jin;Zhuoyi Huang;Lun Yu Li

文献摘要

相似文献

视频语言模型(VLM)是在互联网上大量但嘈杂的视频文本对上预先训练的大型模型,通过其卓越的泛化和开放词汇能力彻底改变了活动识别。虽然复杂的人类活动通常是分层和组合的,但用于评估VLM的大多数现有任务仅关注高级视频理解,因此难以准确评估和解释VLM理解复杂和细粒度人类活动的能力。受最近提出的MOMA框架的启发,我们将活动图定义为人类活动的单一通用表示,包括活动,子活动和原子动作级别的视频理解。我们将活动解析重新定义为活动图生成的首要任务,需要理解所有三个级别的人类活动。为了便于评估活动解析模型,我们引入了MOMA-LRG(多对象多角色细化图),这是一个复杂人类活动的大型数据集,带有活动图注释,可以很容易地转换为自然语言句子。最后,我们提出了一个模型不可知和轻量级的方法来适应和评估VLMs,将结构化的知识从活动图到VLMs,解决语言和图形模型的个别限制。我们在活动解析和少量视频分类方面表现出强大的性能,我们的框架旨在促进未来视频,图形和语言联合建模的研究。
Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional, most existing tasks for evaluating VLMs focus only on high-level video understanding, making it difficult to accurately assess and interpret the ability of VLMs to understand complex and fine-grained human activities. Inspired by the recently proposed MOMA framework, we define activity graphs as a single universal representation of human activities that encompasses video understanding at the activity, subactivity, and atomic action level. We redefine activity parsing as the overarching task of activity graph generation, requiring understanding human activities across all three levels. To facilitate the evaluation of models on activity parsing, we introduce MOMA-LRG (Multi-Object Multi-Actor Language-Refined Graphs), a large dataset of complex human activities with activity graph annotations that can be readily transformed into natural language sentences. Lastly, we present a model-agnostic and lightweight approach to adapting and evaluating VLMs by incorporating structured knowledge from activity graphs into VLMs, addressing the individual limitations of language and graphical models. We demonstrate strong performance on activity parsing and few-shot video classification, and our framework is intended to foster future research in the joint modeling of videos, graphs, and language.