COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality

COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality
复制标题

DOI:
10.1007/978-3-031-19833-5_15
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
Honglu Zhou;Asim Kadav;Aviv Shamsian;Shijie Geng;Farley Lai;Long Zhao;Tingxi Liu;M. Kapadia
Honglu Zhou;Asim Kadav;Aviv Shamsian;Shijie Geng;Farley Lai;Long Zhao;Tingxi Liu;M. Kapadia
中科院分区:
其他
文献类型:
--
作者:
Honglu Zhou;Asim Kadav;Aviv Shamsian;Shijie Geng;Farley Lai;Long Zhao;Tingxi Liu;M. Kapadia

文献摘要

相似文献

群体活动识别检测由一组参与者共同执行的活动,这需要参与者和对象的组合推理。我们通过将视频建模为表示视频中的多尺度语义概念的令牌来完成任务。我们提出了一个基于多尺度Transformer的架构COMPOSESER,它在每个尺度上对令牌进行基于注意力的推理,并学习组合的群体活动。此外,先前的作品遭受场景偏见与隐私和道德问题。我们只使用关键点模态,这减少了场景偏差,并防止获取可能包含用户隐私或偏见信息的详细视觉数据。我们通过聚类中间尺度表示,同时保持一致的规模之间的集群分配,提高了多尺度表示在COMPOSER。最后,我们使用辅助预测和针对关键点信号量身定制的数据增强等技术来辅助模型训练。我们在两个广泛使用的数据集(Volleyball和Collective Activity)上展示了模型的强度和可解释性。COMPOSER仅通过关键点模态实现了改进(代码可在https://github.com/hongluzhou/composer上获得)。
Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by modeling the video as tokens that represent the multi-scale semantic concepts in the video. We proposeCOMPOSER, a Multiscale Transformer based architecture that performs attention-basedreasoningover tokens at each scale and learns group activitycompositionally. In addition, prior works suffer from scene biases with privacy and ethical concerns. We only use the keypoint modality which reduces scene biases and prevents acquiring detailed visual data that may contain private or biased information of users. We improve the multiscale representations inCOMPOSERby clustering the intermediate scale representations, while maintaining consistent cluster assignments between scales. Finally, we use techniques such as auxiliary prediction and data augmentations tailored to the keypoint signals to aid model training. We demonstrate the model’s strength and interpretability on two widely-used datasets (Volleyball and Collective Activity).COMPOSERachieves up toimprovement with just the keypoint modality (Code is available at https://github.com/hongluzhou/composer.).