Active Learning for Video Classification with Frame Level Queries

Active Learning for Video Classification with Frame Level Queries
复制标题

DOI:
10.1109/ijcnn54540.2023.10191348
复制
发表时间:
2023-06
期刊:
2023 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
D. Goswami;Shayok Chakraborty
D. Goswami;Shayok Chakraborty
中科院分区:
其他
文献类型:
--
作者:
D. Goswami;Shayok Chakraborty

文献摘要

相似文献

深度学习算法推动了计算机视觉研究的边界,并在各种应用中描绘了值得称赞的性能。然而,训练一个鲁棒的深度神经网络需要大量的标记训练数据,这些数据的获取需要耗费大量的时间和人力。对于像视频分类这样的应用程序,这个问题甚至更严重,在这种应用程序中,人类注释者必须端到端观看整个视频才能提供标签。主动学习算法从大量未标记的数据中自动识别最有信息的样本;这极大地减少了诱导机器学习模型的人工注释工作,因为只有少数由算法识别的样本需要手动标记。在本文中,我们提出了一种新的视频分类主动学习框架,目的是进一步减少人类注释者的标注负担。我们的框架识别了一批范例视频,以及每个视频的一组信息帧;人类注释者只需要检查帧并为每个视频提供标签。这比观看完整的视频来想出一个标签要少得多的手工工作。我们制定了一个基于不确定性和多样性的标准来识别信息视频,并利用代表性采样技术从每个视频中提取一组示例帧。据我们所知,这是第一个为视频分类开发主动学习框架的研究成果,注释者只需要检查几个帧就可以产生标签,而不是观看端到端的视频。我们广泛的经验分析证实了我们的方法在视频分类等应用程序中大大减少人工注释工作的潜力,在这些应用程序中,注释单个数据实例可能非常繁琐。
Deep learning algorithms have pushed the boundaries of computer vision research and have depicted commendable performance in a variety of applications. However, training a robust deep neural network necessitates a large amount of labeled training data, acquiring which involves significant time and human effort. This problem is even more serious for an application like video classification, where a human annotator has to watch an entire video end-to-end to furnish a label. Active learning algorithms automatically identify the most informative samples from large amounts of unlabeled data; this tremendously reduces the human annotation effort in inducing a machine learning model, as only the few samples that are identified by the algorithm, need to be labeled manually. In this paper, we propose a novel active learning framework for video classification, with the goal of further reducing the labeling onus on the human annotators. Our framework identifies a batch of exemplar videos, together with a set of informative frames for each video; the human annotator needs to merely review the frames and provide a label for each video. This involves much less manual work than watching the complete video to come up with a label. We formulate a criterion based on uncertainty and diversity to identify the informative videos and exploit representative sampling techniques to extract a set of exemplar frames from each video. To the best of our knowledge, this is the first research effort to develop an active learning framework for video classification, where the annotators need to inspect only a few frames to produce a label, rather than watching the end-to-end video. Our extensive empirical analyses corroborate the potential of our method to substantially reduce human annotation effort in applications like video classification, where annotating a single data instance can be extremely tedious.