Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks

Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
复制标题

使用正则化深度神经网络利用视频分类中的特征和类关系

DOI:
10.1109/tpami.2017.2670560
复制
发表时间:
2018-02-01
影响因子:
23.6
通讯作者:
Chang, Shih-Fu
Chang, Shih-Fu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jiang, Yu-Gang;Wu, Zuxuan;Chang, Shih-Fu

文献摘要

被引文献

相似文献

在本文中,我们研究了具有挑战性的问题,根据高层次的语义,如存在一个特定的人类行为或一个复杂的事件的视频分类。虽然近年来已经投入了大量的努力,但大多数现有的作品使用简单的融合策略将多个视频特征结合起来,而忽略了类间语义关系的利用。本文提出了一种新的统一的框架,联合利用的特征关系和类关系,以提高分类性能。具体而言,通过在深度神经网络(DNN)的学习过程中施加正则化来估计和利用这两种类型的关系。通过使DNN具有更好的利用特征和类关系的能力,所提出的正则化DNN(rDNN)更适合于建模视频语义。我们表明,rDNN产生更好的性能超过几个国家的最先进的方法。竞争结果报告了著名的好莱坞2和哥伦比亚消费者视频基准。此外,为了促进未来大规模视频分类的研究,我们收集并发布了一个新的基准数据集,称为FCVID,其中包含91,223个互联网视频和239个手动注释的类别。
In this paper, we study the challenging problem of categorizing videos according to high-level semantics such as the existence of a particular human action or a complex event. Although extensive efforts have been devoted in recent years, most existing works combined multiple video features using simple fusion strategies and neglected the utilization of inter-class semantic relationships. This paper proposes a novel unified framework that jointly exploits the feature relationships and the class relationships for improved categorization performance. Specifically, these two types of relationships are estimated and utilized by imposing regularizations in the learning process of a deep neural network (DNN). Through arming the DNN with better capability of harnessing both the feature and the class relationships, the proposed regularized DNN (rDNN) is more suitable for modeling video semantics. We show that rDNN produces better performance over several state-of-the-art approaches. Competitive results are reported on the well-known Hollywood2 and Columbia Consumer Video benchmarks. In addition, to stimulate future research on large scale video categorization, we collect and release a new benchmark dataset, called FCVID, which contains 91,223 Internet videos and 239 manually annotated categories.