An Adaptive Semisupervised Feature Analysis for Video Semantic Recognition

An Adaptive Semisupervised Feature Analysis for Video Semantic Recognition
复制标题

视频语义识别的自适应半监督特征分析

DOI:
10.1109/tcyb.2017.2647904
复制
发表时间:
2018-02-01
影响因子:
11.8
通讯作者:
Zheng, Qinghua
Zheng, Qinghua
中科院分区:
计算机科学1区
文献类型:
--
作者:
Luo, Minnan;Chang, Xiaojun;Zheng, Qinghua

文献摘要

被引文献

相似文献

视频语义识别通常会受到维数灾难和缺乏足够高质量标记实例的影响,因此半监督特征选择以其高效性和可理解性受到越来越多的关注。大多数先前的方法假设具有近距离(邻居)的视频具有相似的标签,并且通过标记和未标记数据的预定图来表征固有的局部结构。然而,除了图的构造的参数调整问题之外,原始特征空间中的亲和度度量通常遭受维数灾难。此外,预定图将其自身与特征选择过程分离,这可能导致视频语义识别的性能下降。在本文中,我们开发了一种新的半监督特征选择方法从一个新的角度。我们模型的主要假设是,具有相似标签的实例应该有更大的概率成为邻居。而不是使用一个预定的相似性图,我们将探索的局部结构的联合特征选择的过程中,以便同时学习最佳的图。此外,利用自适应损失函数来衡量标签适应度,这显着提高了模型的鲁棒性,视频与小或大量的损失。提出了一种有效的交替优化算法来解决该问题,并从理论上分析了算法的收敛性和计算复杂度。最后,在基准数据集上的大量实验结果表明了该方法在视频语义识别相关任务上的有效性和优越性。
Video semantic recognition usually suffers from the curse of dimensionality and the absence of enough high-quality labeled instances, thus semisupervised feature selection gains increasing attentions for its efficiency and comprehensibility. Most of the previous methods assume that videos with close distance (neighbors) have similar labels and characterize the intrinsic local structure through a predetermined graph of both labeled and unlabeled data. However, besides the parameter tuning problem underlying the construction of the graph, the affinity measurement in the original feature space usually suffers from the curse of dimensionality. Additionally, the predetermined graph separates itself from the procedure of feature selection, which might lead to downgraded performance for video semantic recognition. In this paper, we exploit a novel semisupervised feature selection method from a new perspective. The primary assumption underlying our model is that the instances with similar labels should have a larger probability of being neighbors. Instead of using a predetermined similarity graph, we incorporate the exploration of the local structure into the procedure of joint feature selection so as to learn the optimal graph simultaneously. Moreover, an adaptive loss function is exploited to measure the label fitness, which significantly enhances model’s robustness to videos with a small or substantial loss. We propose an efficient alternating optimization algorithm to solve the proposed challenging problem, together with analyses on its convergence and computational complexity in theory. Finally, extensive experimental results on benchmark datasets illustrate the effectiveness and superiority of the proposed approach on video semantic recognition related tasks.