Recognition of Complex Events: Exploiting Temporal Dynamics between Underlying Concepts

Recognition of Complex Events: Exploiting Temporal Dynamics between Underlying Concepts
复制标题

DOI:
10.1109/cvpr.2014.287
复制
发表时间:
2014-06
期刊:
2014 IEEE Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Subhabrata Bhattacharya;Mahdi M. Kalayeh;R. Sukthankar;M. Shah
Subhabrata Bhattacharya;Mahdi M. Kalayeh;R. Sukthankar;M. Shah
中科院分区:
其他
文献类型:
--
作者:
Subhabrata Bhattacharya;Mahdi M. Kalayeh;R. Sukthankar;M. Shah

文献摘要

被引文献

相似文献

虽然基于特征包的方法在低级别动作分类方面表现出色,但它们不适合识别视频中的复杂事件,其中基于概念的时间表示目前占主导地位。本文提出了一种新的表示,捕捉时间动态的窗口中级概念检测器,以提高复杂事件的识别。我们首先将每个视频表示为有序的向量时间序列,其中每个时间步由从预训练的概念检测器的级联置信度形成的向量组成。我们假设,从同一事件类的不同实例的时间序列的动态,简单的线性动态系统(LDS)模型所捕获的,很可能是相似的,即使实例不同的低层次的视觉功能。我们提出了一个由两部分组成的表示融合:(1)块汉克尔矩阵的奇异值分解(SSID-S)和(2)从相应的特征动力学矩阵计算的谐波签名(HS)。所提出的方法提供了几个替代方法的好处:我们的方法是直接实现,直接采用现有的概念检测器,可以插入到线性分类框架。标准数据集,如NIST的TRECVID多媒体事件检测任务的结果表明,所提出的方法提高了准确性。
While approaches based on bags of features excel at low-level action classification, they are ill-suited for recognizing complex events in video, where concept-based temporal representations currently dominate. This paper proposes a novel representation that captures the temporal dynamics of windowed mid-level concept detectors in order to improve complex event recognition. We first express each video as an ordered vector time series, where each time step consists of the vector formed from the concatenated confidences of the pre-trained concept detectors. We hypothesize that the dynamics of time series for different instances from the same event class, as captured by simple linear dynamical system (LDS) models, are likely to be similar even if the instances differ in terms of low-level visual features. We propose a two-part representation composed of fusing: (1) a singular value decomposition of block Hankel matrices (SSID-S) and (2) a harmonic signature (HS) computed from the corresponding eigen-dynamics matrix. The proposed method offers several benefits over alternate approaches: our approach is straightforward to implement, directly employs existing concept detectors and can be plugged into linear classification frameworks. Results on standard datasets such as NIST's TRECVID Multimedia Event Detection task demonstrate the improved accuracy of the proposed method.