Mining Event Definitions from Queries for Video Retrieval on the Internet

Mining Event Definitions from Queries for Video Retrieval on the Internet
复制标题

DOI:
10.1109/icdmw.2009.70
复制
发表时间:
2009-12
期刊:
2009 IEEE International Conference on Data Mining Workshops
影响因子:
--
通讯作者:
Kimiaki Shirahama;C. Sugihara;Kana Matsumura;Yuta Matsuoka;K. Uehara
Kimiaki Shirahama;C. Sugihara;Kana Matsumura;Yuta Matsuoka;K. Uehara
中科院分区:
其他
文献类型:
--
作者:
Kimiaki Shirahama;C. Sugihara;Kana Matsumura;Yuta Matsuoka;K. Uehara

文献摘要

相似文献

由于互联网上的视频数量巨大且不断增加,因此不可能对这些视频中的事件进行预索引。因此,我们从作为查询提供的示例视频中提取每个事件的定义。但是,与正面例子不同的是,手动提供各种负面例子是不切实际的。因此,我们使用“部分监督学习”,其中事件的定义是从积极的和未标记的示例中提取的。具体地,首先根据正例和未标记例之间的相似性来选择反例。在这里,为了适当地计算相似性,我们使用了表示基于事件中对象的典型布局的相关特征的“视频掩码”。然后,我们从正例和反例中提取事件定义。在这个过程中,我们认为由于不同的摄像技术和物体移动,事件的镜头包含了显著不同的特征。为了涵盖如此大的变化特征,我们使用粗糙集理论来提取事件的多个定义。在TRECVID 2008视频采集上的实验结果验证了该方法的有效性。
Since the amount of videos on the internet is huge and continuously increases, it is impossible to pre-index events in these videos. Thus, we extract the definition of each event from example videos provided as a query. But, different from positive examples, it is impractical to manually provide a variety of negative examples. Hence, we use "partially supervised learning'' where the definition of the event is extracted from positive and unlabeled examples. Specifically, negative examples are firstly selected based on similarities between positive and unlabeled examples. Here, to appropriately calculate similarities, we use a ‘‘video mask'' which represent relevant features based on a typical layout of objects in the event. Then, we extract the event definition from positive and negative examples. In this process, we consider that shots of the event contain significantly different features due to various camera techniques and object movements. In order to cover such a large variation of features, we use "rough set theory'' to extract multiple definitions of the event. Experimental results on TRECVID 2008 video collection validate the effectiveness of our method.