SpotEM: Efficient Video Search for Episodic Memory

SpotEM: Efficient Video Search for Episodic Memory
复制标题

DOI:
10.48550/arxiv.2306.15850
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Santhosh K. Ramakrishnan;Ziad Al-Halah;K. Grauman
Santhosh K. Ramakrishnan;Ziad Al-Halah;K. Grauman
中科院分区:
其他
文献类型:
--
作者:
Santhosh K. Ramakrishnan;Ziad Al-Halah;K. Grauman

文献摘要

相似文献

情景记忆(EM)的目标是搜索长的以自我为中心的视频来回答自然语言查询(例如,“我把钱包放在哪儿了?").现有的EM方法穷尽地提取昂贵的固定长度的剪辑特征,以在视频中的任何地方寻找答案,这对于持续数小时甚至数天的长可穿戴摄像头视频是不可行的。我们提出了SpotEM,一种方法,以实现效率为一个给定的EM方法,同时保持良好的精度。SpotEM由三个关键思想组成:1)一个新的剪辑选择器,它学习识别有希望的视频区域,以根据语言查询进行搜索; 2)一组低成本的语义索引功能,可以捕获房间,对象和交互的上下文,建议去哪里看;以及3)蒸馏损失,其解决由剪辑选择器和EM模型的端到端联合训练引起的优化问题。我们对来自Ego 4D EM自然语言测试基准和三种不同EM模型的200多小时视频的实验证明了我们方法的有效性:仅计算10% - 25%的剪辑特征,我们保留了84% - 97%的原始EM模型的准确性。项目页面:https://vision.cs.utexas.edu/projects/spotem
The goal in episodic memory (EM) is to search a long egocentric video to answer a natural language query (e.g.,"where did I leave my purse?"). Existing EM methods exhaustively extract expensive fixed-length clip features to look everywhere in the video for the answer, which is infeasible for long wearable-camera videos that span hours or even days. We propose SpotEM, an approach to achieve efficiency for a given EM method while maintaining good accuracy. SpotEM consists of three key ideas: 1) a novel clip selector that learns to identify promising video regions to search conditioned on the language query; 2) a set of low-cost semantic indexing features that capture the context of rooms, objects, and interactions that suggest where to look; and 3) distillation losses that address the optimization issues arising from end-to-end joint training of the clip selector and EM model. Our experiments on 200+ hours of video from the Ego4D EM Natural Language Queries benchmark and three different EM models demonstrate the effectiveness of our approach: computing only 10% - 25% of the clip features, we preserve 84% - 97% of the original EM model's accuracy. Project page: https://vision.cs.utexas.edu/projects/spotem