Weakly supervised detection of video events using hidden conditional random fields

Weakly supervised detection of video events using hidden conditional random fields
复制标题

DOI:
10.1007/s13735-014-0068-6
复制
发表时间:
2015-03
影响因子:
5.6
通讯作者:
Kimiaki Shirahama;M. Grzegorzek;K. Uehara
Kimiaki Shirahama;M. Grzegorzek;K. Uehara
中科院分区:
计算机科学4区
文献类型:
--
作者:
Kimiaki Shirahama;M. Grzegorzek;K. Uehara

文献摘要

相似文献

多媒体事件检测(MED)是识别发生特定事件的视频的任务。本文讨论了MED中的两个问题:弱监督设置和事件结构不清晰。第一种情况表明,由于镜头与事件的关联是费力的,并且会引起注释员的主观性,因此对于事件是否包含,培训视频被松散地注释。目前尚不清楚哪些镜头与事件相关或无关。第二个问题是,由于相机和编辑技术的随意性,很难提前假设事件结构。为了解决这些问题,我们提出了一种使用隐藏条件随机场(HCRF)的方法,它是一种具有一组隐藏状态的概率区分分类器。我们认为,弱监督设置可以使用隐藏状态作为中间层来区分与事件相关和不相关的镜头。此外,事件的不清楚结构可以通过每个隐藏状态的特征及其与其他状态的关系来暴露。基于上述思想,我们对隐含状态及其关系进行了优化,以区分包含该事件的训练视频和其他视频。此外,为了充分开发HCRF的潜力,我们建立了训练视频准备、参数初始化和多个HCRF的融合的方法。在TRECVID视频数据上的实验结果验证了该方法的有效性。
Multimedia Event Detection(MED) is the task to identify videos in which a certain event occurs. This paper addresses two problems in MED:weakly supervised settingandunclear event structure. The first indicates that since associations of shots with the event are laborious and incur annotator’s subjectivity, training videos are loosely annotated as to whether the event is contained or not. It is unknown which shots are relevant or irrelevant to the event. The second problem is the difficulty of assuming the event structure in advance, due to arbitrary camera and editing techniques. To tackle these problems, we propose a method using aHidden Conditional Random Field(HCRF) which is a probabilistic discriminative classifier with a set of hidden states. We consider that the weakly supervised setting can be handled using hidden states as the intermediate layer to discriminate between relevant and irrelevant shots to the event. In addition, an unclear structure of the event can be exposed by features of each hidden state and its relation to the other states. Based on the above idea, we optimise hidden states and their relation so as to distinguish training videos containing the event from the others. Also, to exploit the full potential of HCRFs, we establish approaches for training video preparation, parameter initialisation and fusion of multiple HCRFs. Experimental results on TRECVID video data validate the effectiveness of our method.